Software Engineer, Core Infrastructure
Build and operate distributed cloud infrastructure and platform services that support product teams globally. The role requires 5+ years of software development experience, strong distributed-systems expertise, and experience with cloud infrastructure, reliability, and observability.
About the job
Responsibilities
- Design, build, and maintain distributed cloud infrastructure and platform services.
- Work on scaling, automation, reliability, and observability of infrastructure services.
- Operate services, debug issues, and support internal customers.
- Participate in roadmap planning and prioritization.
Requirements
- 5+ years of professional experience in a software development role.
- Experience building, deploying, and managing infrastructure on a major cloud provider.
- Strong engineering background building platform services and/or distributed systems at scale, with a solid grasp of underlying operating system primitives.
- Experience developing, maintaining, and debugging distributed systems, including diagnosing low-level resource constraints.
- Experience with operational excellence and modern observability practices, including distributed tracing, structured logging, and system-level metrics.
Nice-to-haves
- Experience with AWS, Azure, Google Cloud, or Oracle Cloud.
- Experience with Go or other systems languages such as Rust, C, or C++.
- Experience with Linux OS internals, performance optimization, and kernel-level troubleshooting.
- Experience working with Kubernetes clusters and low-level container mechanics.
- Experience in networking and traffic systems at scale.
- Experience handling critical incidents for production systems.
Skills
Distributed Systems, Cloud Infrastructure, Platform Services, AWS, Azure, GCP, Oracle Cloud, Go, Rust, C++, Linux, Kubernetes, Observability, Distributed Tracing, Networking
Similar jobs
DevOps / SRE jobsOwn large-scale ClickHouse cluster upgrades and production operations while building tooling that improves release safety and automation. The role requires 5+ years operating stateful distributed systems, cloud and Kubernetes experience, strong debugging skills, and Go development experience.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.
Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.
Automate, manage, and optimize large-scale ClickHouse clusters handling trillions of events and 100+ PB data. Build provisioning systems with Terraform, Ansible, Kubernetes; focus on performance, scaling, and bleeding-edge features.