Senior Cloud Engineer - Product Metrics
Senior Cloud Engineer responsible for designing, operating, and improving petabyte-scale Product Metrics systems built with Go, Kubernetes, and ClickHouse. The role requires 5+ years of experience with scalable distributed systems and strong reliability, performance, and production debugging skills.
About the job
Responsibilities
- Help define the Product Metrics team roadmap.
- Design, build, operate, and maintain business-critical, petabyte-scale distributed systems.
- Own system performance, reliability, availability, and cost efficiency.
- Deliver and iteratively improve product features in collaboration with the team.
- Mentor teammates, participate in design discussions, and collaborate across engineering teams.
- Participate in an on-call rotation and take ownership of operated services.
Requirements
- 5+ years of relevant software development experience building and operating scalable, fault-tolerant, distributed systems.
- 2+ years of software application development experience using Go.
- Experience with a major cloud service provider such as AWS, Google Cloud, or Azure.
- Experience storing, shipping, and retrieving large data volumes efficiently using technologies such as ClickHouse.
- Experience with Kubernetes, Helm, Argo CD, Temporal, and infrastructure-as-code tools such as Terraform.
- Strong production debugging and problem-solving skills.
- Excellent communication and collaboration skills in a fully remote environment.
Nice-to-haves
- Additional experience with ClickHouse.
- Experience writing Kubernetes operators or controllers.
- Experience with Kafka streaming technology.
Benefits
- Flexible work environment and remote-friendly work culture.
- Employer healthcare contributions.
- Company stock options.
- Flexible time off in the United States and generous entitlement in other countries.
- USD $500 home-office setup allowance for remote employees.
- Opportunities to attend company-wide offsites.
Skills
Go, Kubernetes, ClickHouse, AWS, GCP, Azure, Helm, Argo Cd, Temporal, Terraform, Kafka, Distributed Systems
Similar jobs
DevOps / SRE jobsDesigns and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.
Designs, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.
Build and operate core platform infrastructure, developer tooling, CI/CD, observability, and cloud reliability systems for a regulated payments platform. Requires 5+ years of infrastructure or backend experience, strong infrastructure-as-code skills, and production cloud expertise.
Senior software engineer building standardized, self-service cloud infrastructure across AWS, Google Cloud, and networking systems. Requires 5+ years of software engineering experience, production cloud infrastructure expertise, and proficiency in Go or Python.
Designs and supports physical IT infrastructure across offices, labs, manufacturing facilities, and data centers, including racks, cabling, power, cooling, documentation, and capacity planning. Requires 5+ years of physical infrastructure engineering experience and strong cross-functional project execution.