Skip to content
IllumioIllumio

Staff Engineer

Staff Engineer building scalable, resilient cloud architecture and distributed systems on AWS/Azure/GCP with Kubernetes for Illumio's Zero Trust cybersecurity platform. Requires strong cloud programming experience, distributed systems expertise, and 7+ years overall experience.

About the job

Your Impact

  • Deliver cloud architecture that solves business problems while balancing architecture and business margins.
  • Collaborate with cross-functional teams, including Product Managers, developers, and DevOps engineers to understand business requirements and architect scalable cloud solutions.
  • Design, deploy, and manage cloud-based architecture using industry best practices, focusing on efficiency, scalability, availability, performance, and security.
  • Evaluate and select appropriate cloud technologies, platforms, and tools, including Kubernetes, to meet the company's needs and drive innovation.
  • Optimize cloud-based systems for high availability, fault tolerance, and disaster recovery.
  • Implement and manage monitoring, logging, and alerting systems to ensure the health and performance of the cloud infrastructure.
  • Identify and resolve performance bottlenecks, security vulnerabilities, and other operational issues in the cloud environment.
  • Stay up-to-date with the latest trends, technologies, and best practices in cloud computing, distributed systems, and cybersecurity.

Your Toolkit

  • Proven experience as a Cloud Architect or similar role, designing and implementing cloud-based solutions in a production environment.
  • AWS / Azure / GCP cloud experience: extensively used one of these platforms at the API/programming level.
  • Experience with networking and security controls is a plus.
  • Experience in one or more programming languages: Java, Python.
  • REST API client experience.
  • CloudFormation, Terraform are nice to have.
  • Working knowledge of Spring or similar ecosystems.
  • Knowledge of Kubernetes and Docker containers is good to have.
  • Gang of Four design patterns.
  • Strong expertise in cloud platforms such as Amazon Web Services (AWS), Microsoft Azure, or Google Cloud Platform (GCP).
  • In-depth knowledge of Kubernetes and containerization technologies.
  • Built highly scalable distributed systems from the ground up.
  • Experience with building highly scalable and resilient cloud services.
  • Excellent problem-solving and troubleshooting skills.
  • Strong communication and collaboration abilities.
  • Ability to work in an agile environment and drive initiatives with clarity and purpose.

Skills

AWS, Azure, GCP, Kubernetes, Docker, Java, Python, Terraform, CloudFormation, REST APIs, Spring

Crusoe

Crusoe

San Francisco, CA
Staff Network Engineer, Operations
$195k+/yrOn-site8+ YOEDevOps / SRE

Own reliability, incident response, observability, and automation for Crusoe Cloud’s global network infrastructure supporting large-scale GPU workloads. The role requires 8+ years of production network engineering experience, expertise in data center and lossless fabrics, Python automation skills, and strong operational leadership.

Shield AI

Shield AI

San Diego, CA

Senior Staff Lead Site Reliability Engineer
$190k+/yrOn-site7+ YOEDevOps / SRE

Leads the establishment and maturation of SRE practices across cloud infrastructure and platform services. This hands-on technical role focuses on reliability targets, observability, incident response, resilience, automation, and mentoring engineering teams.

OpenSea

OpenSea

United States

Staff Platform Engineer
$190k+/yrRemote7+ YOEDevOps / SRE

Build and operate scalable platform services, infrastructure, and developer tooling that enable reliable product delivery. The role requires 7+ years of software engineering experience, JVM expertise, distributed-systems experience, and strong platform, cloud, CI/CD, and observability skills.

Komodo Health

Komodo Health

United States

Staff Infrastructure Engineer
$187k+/yrRemote8+ YOEDevOps / SRE

Leads architecture, ownership, modernization, and operation of Komodo Health’s AWS and Kubernetes infrastructure and shared services. The role requires 8+ years of infrastructure experience, deep Terraform and Kubernetes expertise, regulated-environment security fluency, and the ability to establish AI-assisted engineering standards.

VGS

VGS

United States
Senior Staff Infrastructure Engineer
$185k+/yrRemote10+ YOEDevOps / SRE

Leads the architecture, automation, observability, and reliability of multi-region AWS infrastructure supporting high-throughput payments. Requires 10+ years of distributed-systems experience and deep expertise in cloud infrastructure, Kubernetes, infrastructure as code, and modern SRE practices.