Skip to content
IllumioIllumio

Sr. Member of Technical Staff - Platform

Designs, builds, and maintains cloud-native platform services using Kubernetes for multi-tenant deployments. Develops microservices in Ruby/Java/Go, manages CI/CD pipelines with GitOps, and ensures production reliability across AWS/Azure. Requires 3-5 years experience and strong Kubernetes expertise.

About the job

Responsibilities

  • Design, build, and maintain core cloud-native platform services utilizing the Kubernetes control plane to enable rapid product development and high-scale deployment in a multi-tenant, distributed architecture.
  • Develop robust microservices using Ruby, Java, or Go, using Kubernetes and data services like PostgreSQL and Redis across Azure and AWS clusters.
  • Own the full development lifecycle including deployment, focusing on defining Kubernetes deployment strategies, ensuring comprehensive observability (Prometheus/Grafana), and implementing best practices for production reliability and critical issue resolution.
  • Enhance and manage multi-cluster CI/CD pipelines using Jenkins and automation, strictly adhering to GitOps principles (e.g., ArgoCD or Flux) for declarative configuration management across hybrid and multi-cloud Kubernetes environments.
  • Foster team excellence through mentorship and rigorous code reviews, driving adoption of Kubernetes security policies, resource optimization, and overall cloud-native best practices.

Requirements

  • Bachelor’s degree in Computer Science or equivalent, with 3–5 years of experience building distributed, scalable software and systems.
  • Strong coding skills in Ruby, Java, or Go, with experience building API-based web services and a desire to deepen expertise in Ruby.
  • Hands-on experience with PostgreSQL, Redis, or similar database and caching technologies.
  • Deep expertise in Kubernetes architecture and operations, with familiarity implementing GitOps workflows (e.g., ArgoCD, Flux).
  • Solid understanding of CI/CD systems (especially Jenkins), and experience working with Azure and AWS at the API/programming level.
  • Demonstrated analytical and debugging skills, with a passion for continuous learning and staying current with technological trends.
  • Excellent communication and collaboration abilities, thriving in team-oriented environments.

Nice-to-Haves

  • Experience with networking and network security controls.
  • Exposure to platform resiliency practices and incident response.
  • Familiarity with observability tools (e.g., Prometheus, Grafana).
  • Experience with infrastructure as code (e.g., Terraform, Pulumi).

Skills

Kubernetes, Ruby, Java, Go, Postgres, Redis, GitOps, Argo CD, Flux, Jenkins, AWS, Azure, Prometheus, Grafana, Terraform

Imply

Imply

United States

Senior Software Engineer
$155k+/yrRemote6+ YOEDevOps / SRE

Build and operate highly available, distributed platform services and cloud infrastructure for petabyte-scale observability products. The role requires 6+ years of experience, strong Java and AWS expertise, Kubernetes and Terraform production experience, and a bachelor’s degree or equivalent.

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.

Gumloop

Gumloop

San Francisco, CA
Senior Infrastructure Engineer
$150k+/yrOn-siteDevOps / SRE

Own and scale infrastructure for agent orchestration, sandboxing, and hosted MCP services. The role requires hands-on Kubernetes, cloud, and infrastructure-as-code experience, along with strong software engineering fundamentals and high ownership.

Axle

Axle

Frederick, MD

IT Operations Technical Lead
$150k+/yrHybrid10+ YOEDevOps / SRE

Leads hybrid cloud and on-premises IT operations, incident management, automation, security hardening, and infrastructure reliability while mentoring systems engineers. Requires extensive Linux administration, ITIL operations, cloud migration, automation, and AI/ML infrastructure experience.

PointClickCare

PointClickCare

United States

Senior Site Reliability Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

Senior Site Reliability Engineer providing technical leadership for scalable operations, automation, monitoring, resiliency, and cloud infrastructure. Requires a bachelor's degree, software development or architecture experience, and hands-on DevOps or systems administration experience.