Skip to content

Senior IT Site Reliability Software Engineer

Builds and operates resilient, observable cloud infrastructure and automation for internal IT services. The role requires 5+ years of production software engineering experience, strong Python skills, infrastructure-as-code expertise, and hands-on cloud and container experience.

About the job

Responsibilities

  • Design and deploy production-grade infrastructure on AWS or Azure using Infrastructure as Code tools such as Terraform or Pulumi.
  • Optimize system performance, architecture, and scaling to maximize uptime and minimize latency for critical IT services.
  • Architect robust CI/CD pipelines using GitHub Actions, including hosted and self-hosted runners.
  • Build infrastructure that enables internal applications to include security, logging, metrics, and alerts by default.
  • Build internal AI plugins and automation scripts to streamline developer workflows and improve operational efficiency.
  • Develop incident management workflows and dashboards to maintain service health.
  • Participate in a shared on-call rotation and lead incident response and technical troubleshooting for production outages.
  • Facilitate blameless post-mortems, identify root causes, and implement permanent preventive engineering solutions.
  • Collaborate with Security, Engineering, and Support teams.

Requirements

  • 5+ years of production-level software engineering experience.
  • Strong proficiency in Python.
  • Expert-level proficiency in Terraform, including modules and state management, or Pulumi.
  • Hands-on experience with AWS, Azure, or GCP.
  • Experience with Kubernetes, Docker, and containerization concepts.
  • Deep understanding of observability pillars: logging, metrics, and tracing.
  • Experience with observability tools such as Datadog, Prometheus, or ELK.
  • Proficiency with distributed systems concepts, including Kafka or messaging queues.
  • Advanced knowledge of GitHub Actions and GitHub Runners.
  • Ability to own ambiguous projects and execute independently with minimal guidance.

Benefits

  • Comprehensive benefits and perks tailored to employees in the relevant region.

Skills

Python, Terraform, Pulumi, AWS, Azure, GCP, Kubernetes, Docker, GitHub Actions, Github Runners, Datadog, Prometheus, Elk, Kafka, Infrastructure As Code

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.

Lightspark

Lightspark

Remote

Senior Production Engineer
$200k+/yrRemote5+ YOEDevOps / SRE

The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.

Clickhouse

Clickhouse

Remote

Senior Cloud Software Engineer - Efficiency Engineering
No salary listedRemote5+ YOEDevOps / SRE

Designs and operates scalable, highly available cloud infrastructure while leading efficiency initiatives across compute, storage, networking, and cost optimization. Requires 5+ years of distributed-systems software development experience and expertise with cloud platforms, infrastructure as code, and Kubernetes.

Clickhouse

Clickhouse

Remote

Senior Cloud Software Engineer - Efficiency Engineering
No salary listedRemote5+ YOEDevOps / SRE

Build and optimize ClickHouse Cloud’s highly available, multi-cloud infrastructure, including automation, distributed systems, networking, security, and cost-efficiency tooling. Requires 5+ years of experience operating scalable systems and expertise in cloud platforms, infrastructure as code, and production engineering.

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.