Skip to content
StripeStripe

Software Engineer, Core Infrastructure

Build and operate distributed cloud infrastructure and platform services that support product teams globally. The role requires 5+ years of software development experience, strong distributed-systems expertise, and experience with cloud infrastructure, reliability, and observability.

About the job

Responsibilities

  • Design, build, and maintain distributed cloud infrastructure and platform services.
  • Work on scaling, automation, reliability, and observability of infrastructure services.
  • Operate services, debug issues, and support internal customers.
  • Participate in roadmap planning and prioritization.

Requirements

  • 5+ years of professional experience in a software development role.
  • Experience building, deploying, and managing infrastructure on a major cloud provider.
  • Strong engineering background building platform services and/or distributed systems at scale, with a solid grasp of underlying operating system primitives.
  • Experience developing, maintaining, and debugging distributed systems, including diagnosing low-level resource constraints.
  • Experience with operational excellence and modern observability practices, including distributed tracing, structured logging, and system-level metrics.

Nice-to-haves

  • Experience with AWS, Azure, Google Cloud, or Oracle Cloud.
  • Experience with Go or other systems languages such as Rust, C, or C++.
  • Experience with Linux OS internals, performance optimization, and kernel-level troubleshooting.
  • Experience working with Kubernetes clusters and low-level container mechanics.
  • Experience in networking and traffic systems at scale.
  • Experience handling critical incidents for production systems.

Skills

Distributed Systems, Cloud Infrastructure, Platform Services, AWS, Azure, GCP, Oracle Cloud, Go, Rust, C++, Linux, Kubernetes, Observability, Distributed Tracing, Networking

Clickhouse

Clickhouse

Singapore
Release Engineer - Data Plane Internal Tooling and Productivity
No salary listedRemote5+ YOEDevOps / SRE

Own large-scale ClickHouse cluster upgrades and production operations while building tooling that improves release safety and automation. The role requires 5+ years operating stateful distributed systems, cloud and Kubernetes experience, strong debugging skills, and Go development experience.

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

Supabase

Supabase

Remote

Platform Engineer - Compute Capacity
No salary listedRemote5+ YOEDevOps / SRE

Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.

Alpaca

Alpaca

Remote

Production Support Engineer
No salary listedRemote4+ YOEDevOps / SRE

Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.

PostHog

PostHog

Remote

ClickHouse Operations Engineer
No salary listedRemoteDevOps / SRE

Automate, manage, and optimize large-scale ClickHouse clusters handling trillions of events and 100+ PB data. Build provisioning systems with Terraform, Ansible, Kubernetes; focus on performance, scaling, and bleeding-edge features.