Build and optimize a next-generation CI platform infrastructure for workflow orchestration, scheduling, caching, and autoscaling to accelerate software development for engineers and AI coding agents at scale. Requires interest in distributed systems and knowledge of Kubernetes, EC2, Golang, and Docker.
162k – 190k/yr
RemoteDevOps / SRE
About the role
Responsibilities
Experiment with and prototype new features in a remote sandbox development environment to enable new capabilities in CI-related workflows.
Debug workflow orchestration queueing delays and measure the performance of remote execution/cache infrastructure to improve correctness, reproducibility, and speed.
Identify root causes of infrastructure-related reasons for test flakiness in CI.
Optimize the speed of autoscaling CI workers to handle bursts in CI workloads from product engineering teams.
Up-level observability and alerting of the CI platform across build systems, flaky test management systems, and the merge queue.
Build an improved canarying mechanism to roll out workflow orchestration changes iteratively and safely.
Design and build key pieces of a new CI platform for workflow orchestration, focusing on scheduling, queueing, backpressure, consistency, caching, fault tolerance, and observability.
Work on scheduling work onto the most cost-efficient compute stack.
Requirements
Bachelor’s and/or Master’s degree, preferably in CS, or equivalent experience.
Interest in working on distributed systems engineering and optimization problems.
General knowledge of most of the following: Kubernetes, EC2, Golang, Docker.
Metrics-driven approach to decision-making, with an interest in delivering quantifiable impact.
Nice-to-Haves
Experience with CI platforms, build systems, or workflow orchestration at scale.
Background in optimizing for AI-scale software development workloads.
IT Systems Engineer builds, manages, and supports Windows/Linux infrastructure, VMware virtualization, and Puppet automation for corporate systems. Requires 3-5 years experience in systems engineering, troubleshooting, scripting, and on-call support in a fast-paced environment.
162k – 226k/yr
On-site3+ YOEDevOps / SRE
Infrastructure Engineer
TailscaleUnited States
Infrastructure Engineer builds and maintains internal engineering services, improves observability, CI/CD pipelines, and cloud infrastructure using tools like Kubernetes and AWS. Requires experience with distributed systems, infrastructure as code, and operating managed services in a remote environment.
163k – 204k/yr
RemoteDevOps / SRE
Software Engineer, Software Update Infrastructure
NuroMountain View, CA
Builds and maintains release and OTA update infrastructure for self-driving vehicle fleets, focusing on cloud-robot connectivity, telemetry, and scalable systems. Requires 5+ years in large-scale distributed systems with strong C++ or Go proficiency.
160k – 241k/yr
On-site5+ YOEDevOps / SRE
Site Reliability Engineer
RetoolSan Francisco, CA +1
Site Reliability Engineer owning reliability, automation, and upgrades across Retool Cloud, BYOC, and self-hosted Kubernetes deployments for enterprise customers. Requires deep AWS, Kubernetes, Terraform, and Postgres experience plus strong automation and customer collaboration skills.
164k – 306k/yr
Hybrid5+ YOEDevOps / SRE
Software Engineer, Core Infrastructure
RetoolSan Francisco, CA
Builds and scales core cloud infrastructure for high availability, performance, and developer productivity. Collaborates across teams to evolve backend architecture, automate workflows, and maintain production systems using Kubernetes, Docker, Postgres, and Node.