Skip to content
ZillizZilliz

Senior Site Reliability Engineer Cloud Platform

Senior SRE focuses on ensuring reliability, availability, and performance of distributed database systems in cloud-native environments. Requires 4+ years experience with Kubernetes, Docker, cloud platforms (AWS/GCP/Azure), IaC tools, and scripting in Python/Go/Java.

About the job

Responsibilities

  • Work at the intersection of development and site reliability, creating SRE tools and systems while supporting existing infrastructure and platforms.
  • Ensure the reliability, availability, and performance of Zilliz’s distributed database systems.
  • Develop and implement strategies for monitoring, incident management, and disaster recovery.
  • Automate system operations and maintenance tasks to improve efficiency and reduce manual intervention.
  • Design and build tools to manage and monitor infrastructure, ensuring scalability and robustness.
  • Collaborate with software engineers to enhance system reliability, scalability, and performance.
  • Maintain and improve the CI/CD pipeline to ensure smooth and rapid deployment of changes.
  • Actively contribute to the Milvus Vector Database open-source community, focusing on improving reliability and operational efficiency.

Requirements

  • 4+ years of experience in site reliability engineering or similar roles with a focus on cloud-native systems.
  • Proficiency in scripting languages such as Python, Go, or Java.
  • Strong knowledge of container orchestration technologies like Kubernetes and Docker.
  • Expertise with cloud platforms such as AWS, GCP, or Azure, and their respective monitoring and management tools.
  • Experience with infrastructure as code tools such as Terraform or Ansible.
  • Familiarity with CI/CD tools such as Jenkins, GitLab CI, or Argo.
  • Proven ability to troubleshoot complex distributed systems and resolve issues promptly.
  • Bachelor’s degree or above in computer science, software engineering, or other relevant disciplines.
  • Ability to thrive in a fast-paced, startup environment and handle multiple projects simultaneously.

Nice-to-Haves

  • Experience with Open Source Milvus Vector Database.

Skills

Python, Go, Java, Kubernetes, Docker, AWS, GCP, Azure, Terraform, Ansible, Jenkins, Gitlab Ci, Argo, Milvus

EliseAI

EliseAI

New York, NY

Senior Platform Operations Engineer
$175k+/yrOn-site5+ YOEDevOps / SRE

Leads cross-functional technical initiatives and builds scalable business operations and customer-facing systems. Requires Python, system design, production engineering experience, and strong stakeholder collaboration; platform, AWS, SaaS, and analytics experience are preferred.

Okta

Okta

Bellevue, WA

Senior Manager, Site Reliability Engineering - Infrastructure Platform
$176k+/yrHybrid6+ YOEDevOps / SRE

Leads Infrastructure Platform and Shared Services teams, overseeing Edge networking, Kubernetes platform, CI/CD, observability, and automation. Requires 6+ years technical leadership, AWS expertise, and strong Kubernetes/Terraform skills.

Temporal

Temporal

United States

Senior Software Engineer, Infrastructure Foundations
$176k+/yrRemote10+ YOEDevOps / SRE

Build and scale reliable cloud infrastructure systems, shape long-term architecture and roadmaps, and drive cross-functional alignment. The role requires 10+ years of coding experience, distributed-systems and concurrency expertise, deep infrastructure experience, and hands-on cloud-provider experience.

Runpod

Runpod

United States

Senior HPC Storage Engineer
$180k+/yrRemote8+ YOEDevOps / SRE

Own the design, scaling, reliability, and automation of a multi-region storage platform supporting AI workloads. The role requires 8+ years of production infrastructure or storage engineering experience, distributed storage expertise, strong Linux and networking knowledge, and production programming skills.

tastytrade

tastytrade

Chicago, IL

Senior Site Reliability Engineer - Linux Systems & Application Observability
$180k+/yrHybrid5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for building fault-tolerant infrastructure, scaling a Nomad-based service fabric, and strengthening observability for critical brokerage systems. The role requires production experience with distributed systems, Linux, networking, instrumentation, on-call operations, and reliability practices.