Skip to content
SocureSocure

Senior Software Engineer - SRE

As a Senior Site Reliability Engineer, you will own the end-to-end reliability and scalability of AWS infrastructure and Kubernetes platforms. This role involves designing, operating, and continuously improving production systems with a strong focus on automation and observability.

About the job

What You’ll Own

  • End-to-end ownership of highly available, scalable AWS infrastructure
  • Design, operation, and continuous improvement of Kubernetes (EKS) platforms
  • Reliability of production systems through strong observability, automation, and SLOs
  • CI/CD systems that enable safe, fast, and repeatable deployments
  • Infrastructure defined and enforced through Terraform and GitOps
  • Incident response, root cause analysis, and long-term remediation
  • Raising operational standards through automation, documentation, and best practices

Technical Requirements:

We’re looking for engineers who have actually built, run, and scaled real production systems in the following areas:

Cloud & Infrastructure

  • Deep AWS expertise - networking, compute, IAM, scaling, security
  • Strong experience managing infrastructure using Terraform at scale

Kubernetes & Platform Engineering

  • Very strong Kubernetes fundamentals (internals, scheduling, networking, storage)
  • Hands-on experience operating Amazon EKS in production environments
  • Experience troubleshooting complex, multi-layer Kubernetes issues

Coding & Automation

  • Ability to write clean, maintainable, production-quality code in: Go/ Python
  • Strong automation mindset — eliminating toil through code

CI/CD & GitOps

  • Proven experience building and operating CI/CD pipelines
  • Hands-on experience with:
    • GitHub (Actions or integrations)
    • ArgoCD and GitOps-based deployment workflows

Observability & Reliability

  • Strong understanding of observability principles: metrics, logs, traces, and alerting
  • Hands-on experience with Datadog or similar tool for:
    • Infrastructure and Kubernetes monitoring
    • Application performance monitoring (APM)
    • Alerting, dashboards, and incident detection
  • Experience defining and using SLIs/SLOs to drive reliability decisions
  • Ability to turn observability data into actionable operational improvements

Skills

AWS, Terraform, Kubernetes, Amazon Eks, Go, Python, CI/CD, GitHub Actions, Argo CD, Datadog

Voltus

Voltus

United States

Senior Software Engineer, Infrastructure
$160k+/yrRemote6+ YOEDevOps / SRE

Senior software engineer responsible for operating and evolving Voltus’s infrastructure platform across AWS, Kubernetes, Nomad, observability, stateful systems, and developer tooling. The role requires 6+ years of engineering experience, deep production Kubernetes and AWS expertise, and strong Go or Python skills.

VSCO

VSCO

San Francisco, CA

Senior Software Engineer, Infrastructure
$165k+/yrHybrid5+ YOEDevOps / SRE

Own and evolve VSCO’s AWS/EKS platform, including infrastructure as code, GitOps, CI/CD, observability, networking, and production reliability. The role requires 5+ years of hands-on infrastructure or SRE experience and strong Kubernetes, Terraform, and AWS expertise.

Imply

Imply

United States

Senior Software Engineer
$155k+/yrRemote6+ YOEDevOps / SRE

Build and operate highly available, distributed platform services and cloud infrastructure for petabyte-scale observability products. The role requires 6+ years of experience, strong Java and AWS expertise, Kubernetes and Terraform production experience, and a bachelor’s degree or equivalent.

Commure

Commure

Mountain View, CA
Senior Software Engineer, Infrastructure
$170k+/yrHybrid6+ YOEDevOps / SRE

Own foundational cloud infrastructure and the internal developer platform supporting Commure’s engineering teams. The role requires 6+ years of infrastructure, platform, or SRE experience and hands-on expertise across Kubernetes, infrastructure as code, GitOps, observability, and cloud environments.

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.