Infrastructure Operations Engineer
Build and operate infrastructure platforms supporting internal and customer-facing workloads, with a focus on reliability, automation, cloud, Kubernetes, storage, and networking. The role requires extensive Linux and AWS experience plus hands-on infrastructure-as-code and systems engineering skills.
About the job
Responsibilities
- Design, build, and roll out platforms and patterns that minimize incidents and enable customer-facing and internal features.
- Deploy updates and improvements for internal and end-customer use cases.
- Collaborate with Infrastructure Engineering, Network Operations, Customer Success, and Software and Platform Development teams.
- Participate in an evenly distributed primary/secondary on-call rotation.
Requirements
- 8+ years working with Linux as a server or hosting platform; Ubuntu experience is a plus.
- 5+ years of experience with AWS.
- 2+ years of experience with Kubernetes and strong container fundamentals.
- 2+ years of experience with Terraform and Ansible.
- 2+ years managing network-attached storage using NFS, Ceph, or other protocols; VAST storage experience is a plus.
- Experience with monitoring systems such as Prometheus and the ELK stack.
- Familiarity with GitOps workflows.
- Software development experience using Python, Go, Bash, or other languages for automation and system/API integration.
- Deep networking fundamentals; data-center networking, 400Gb Ethernet, and InfiniBand experience are pluses.
- Experience building and delivering complex systems.
- Ability to navigate tradeoffs among design, risk, cost, and outcomes.
- Comfort navigating ambiguity.
- Strong written and oral communication.
Nice-to-Haves
- Bare-metal hardware troubleshooting and provisioning, particularly Dell hardware.
- Experience with GPU servers in bare-metal or virtualized environments.
- Deep experience with network switches, routers, and firewalls, particularly SONiC switches, Palo Alto firewalls, and Juniper Networks equipment.
- Experience with VAST storage systems.
Compensation and Benefits
- Anticipated annual base salary: SGD 165,000–205,000.
- Discretionary bonus and meaningful equity component.
- Comprehensive medical, dental, and vision coverage for employees and eligible dependents.
- RSUs.
- Flexible time off, company holidays, and floating holidays.
- Two-week company-wide winter break.
- Paid parental and family leave.
- Annual learning and development allowance.
- Wellness and work-from-home stipends.
- Four weeks of paid sabbatical leave after four years of service.
- Flexible schedules and a hybrid work model for office-based teams.
- Complimentary meals at office hubs.
- Benefits may vary by location, team, and role.
Skills
Linux, Ubuntu, AWS, Kubernetes, Terraform, Ansible, Nfs, Ceph, Prometheus, Elk Stack, GitOps, Python, Go, Bash, Networking
Similar jobs
DevOps / SRE jobsThe Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Designs and operates scalable, highly available cloud infrastructure while leading efficiency initiatives across compute, storage, networking, and cost optimization. Requires 5+ years of distributed-systems software development experience and expertise with cloud platforms, infrastructure as code, and Kubernetes.
Build and optimize ClickHouse Cloud’s highly available, multi-cloud infrastructure, including automation, distributed systems, networking, security, and cost-efficiency tooling. Requires 5+ years of experience operating scalable systems and expertise in cloud platforms, infrastructure as code, and production engineering.
Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.