Build and scale reliable cloud infrastructure systems, shape long-term architecture and roadmaps, and drive cross-functional alignment. The role requires 10+ years of coding experience, distributed-systems and concurrency expertise, deep infrastructure experience, and hands-on cloud-provider experience.
176k – 238k/yr
Remote10+ YOEDevOps / SRE
About the role
Responsibilities
Invent
Build systems that optimize cloud use and scale for expansion.
Contribute to system architecture and execution.
Co-author roadmaps establishing the vision for infrastructure scalability, reliability, and engineering velocity.
Maintain and foster a culture of engineering and operational excellence.
Own
Develop effective partnerships across engineering and product teams.
Collaborate across teams to keep roadmaps aligned.
Perform design reviews for multiple projects, focusing on infrastructure scalability and reliability.
Partner with Security to build secure-by-default systems.
Engage with stakeholders and customers to understand their requirements and enable their business.
Learn
Investigate how to best leverage Temporal’s software to power infrastructure at scale.
Understand customer needs that drive infrastructure design decisions.
Collaborate
Share design principles for building reliable systems at scale.
Help others improve while continuing to grow professionally.
Requirements
Experience contributing to complex, cross-team engineering efforts focused on cloud, compute, networking, and storage infrastructure.
Excellent collaboration and communication skills.
Ability to drive alignment across an organization and contribute to long-term roadmaps.
Demonstrated ability to make system-level tradeoffs.
At least 10 years of coding experience using Go, Java, or another applicable language, including experience writing concurrent code.
Experience designing distributed systems and using concurrency primitives.
Deep experience in at least one infrastructure domain and familiarity with adjacent domains.
Hands-on experience with one or more cloud providers, such as AWS, GCP, or Azure.
Compensation and Benefits
Estimated pay range: $176,000–$237,600.
Eligible to participate in Temporal’s equity plan.
Unlimited PTO, 12 holidays, and 2 floating holidays.
100% premium coverage for medical, dental, and vision insurance.
AD&D, short- and long-term disability, and life insurance.
Empower 401(k) plan.
Additional perks for learning and development, lifestyle spending, home office setup, professional memberships, work-from-home meals, internet stipend, and more.
Occasional travel may be required for company events, team offsites, and other in-person opportunities.
Leads end-to-end development of scalable distributed systems for infrastructure observability, owns production issues, and collaborates on designs. Requires expertise in Go, Kubernetes, SQL, cloud providers, and observability tools like Clickhouse and Prometheus.
176k – 238k/yrRemoteDevOps / SRE
Sr. Software Engineer, Platform Engineering
ChronographUnited States
Platform engineer responsible for designing, building, and operating cloud infrastructure on AWS and Kubernetes. Focus on developer tooling, automation with AI, performance tuning, security, observability, and on-call support for a SaaS platform serving private capital markets. Requires 5+ years in DevOps/SRE/Platform roles.
175k – 215k/yrRemote5+ YOEDevOps / SRE
Senior Software Engineer, Cloud Platform
ZillizRedwood City, CA
Build and operate the cloud platform powering Zilliz Cloud and Vector Lakebase across multi-cloud environments, integrating control plane, scheduling, and database runtime for scalable AI workloads. Requires 3+ years building production systems, strong Kubernetes and cloud experience, and a bachelor's degree or equivalent.
175k – 225k/yrHybrid3+ YOEDevOps / SRE
Senior Site Reliability Engineer Cloud Platform
ZillizRedwood City, CA
Senior SRE focuses on ensuring reliability, availability, and performance of distributed database systems in cloud-native environments. Requires 4+ years experience with Kubernetes, Docker, cloud platforms (AWS/GCP/Azure), IaC tools, and scripting in Python/Go/Java.
175k – 225k/yrHybrid4+ YOEDevOps / SRE
Senior Site Reliability Engineer
FivetranOakland, CA
Senior SRE responsible for production infrastructure reliability, incident response, deployment automation, and scaling SaaS systems on Kubernetes and major cloud platforms.