Skip to content
PigmentPigment

Senior SRE

The Senior SRE will design, automate, and operate highly available infrastructure for a real-time SaaS platform handling massive datasets. The role combines reliability engineering, observability, incident response, security collaboration, and software development.

About the job

Responsibilities

  • Design and implement scalable infrastructure for a real-time platform processing and synchronizing large datasets.
  • Automate infrastructure scaling in and out according to platform usage.
  • Ensure high availability and redundancy of the platform.
  • Monitor platform performance and correctness, and promote observability best practices across the engineering team.
  • Participate in incident response.
  • Support geographical expansion as the company grows and serves clients overseas.
  • Work with the security team on infrastructure and development pipeline initiatives, including code repositories and credentials management.
  • Identify inefficiencies in development practices and pipelines, and drive improvements and automation across the software engineering team.
  • Contribute to software development activities and collaborate closely with software engineers.

Requirements

  • Experience as a software engineer and in DevOps or SRE.
  • Experience with a container orchestration platform; Kubernetes is a plus.
  • Experience with a public cloud provider; Google Cloud Platform is a plus.
  • Proven software development experience with languages such as C#, Java, C++, Golang, Rust, JavaScript, Python, or Ruby.
  • Experience with observability tools such as Datadog, Prometheus, ELK, or Jaeger.
  • Strong teamwork and problem-solving skills.
  • Humility and willingness to grow.
  • Fluency in English.

Technical Stack

  • Kubernetes, GKE, and Google Cloud Platform
  • Terraform
  • PostgreSQL, Cloud SQL, CloudNativePG, SingleStore, and Elasticsearch
  • RabbitMQ
  • Temporal.io
  • ArgoCD, Istio, Vault, GitHub, and Docker
  • C# ASP.NET Core 8 and Golang
  • React, TypeScript, Jest, Cypress, and Vite

Skills

Kubernetes, GCP, Terraform, Postgres, Singlestore, Elasticsearch, RabbitMQ, Temporal.Io, Argo CD, Istio, Vault, Docker, Go, Prometheus, React

Hudl

Hudl

London, United Kingdom

Senior Engineer - Platform
£66k+/yrRemote5+ YOEDevOps / SRE

Senior Platform Engineer responsible for architecting scalable, secure infrastructure and improving reliability, observability, and production operations. The role requires strong AWS, Infrastructure as Code, and Kubernetes experience, along with technical leadership and mentoring skills.

Mozilla

Mozilla

Canada

Senior Software Engineer, Cloud Engineering
CA$95k+/yrRemote5+ YOEDevOps / SRE

Build and mature Mozilla’s internal developer infrastructure platform, including CI/CD, observability, Kubernetes optimization, environment bootstrapping, and cost optimization. The role requires 5+ years of software engineering experience, cloud-native expertise, and strong technical leadership.

Muck Rack

Muck Rack

Bulgaria
Senior Software Engineer, DevOps
€95k+/yrRemote5+ YOEDevOps / SRE

Senior DevOps Engineer responsible for building and operating Kubernetes-based infrastructure, AWS cloud systems, deployment workflows, and observability for reliable services at scale. Requires 5+ years of DevOps or platform engineering experience and strong production Kubernetes expertise.

Applied Intuition

Applied Intuition

London, United Kingdom

Senior Cloud Infrastructure Engineer
£100k+/yrOn-site5+ YOEDevOps / SRE

The Senior Cloud Infrastructure Engineer will design and operate secure cloud, on-premise, and air-gapped infrastructure for UK defence customers. The role requires 5+ years of production infrastructure experience, cloud and IaC expertise, Kubernetes and containerization skills, and eligibility for UK Security Clearance.

Okta

Okta

Bellevue, WA
Senior Site Reliability Engineer
$147k+/yrHybrid5+ YOEDevOps / SRE

The Senior Site Reliability Engineer will build and operate secure, highly available infrastructure and Snowflake data tooling for large-scale SaaS systems. The role emphasizes automation, Kubernetes, Terraform, CI/CD, incident response, and collaboration with development, data science, and security teams.