Senior Infrastructure Engineer
Builds and owns foundational infrastructure for globally distributed systems, implements SRE objectives in Golang, manages Kubernetes clusters, and leads incident response. Requires expertise in software engineering, systems administration, and multi-region operations.
About the job
What You'll Do
- Build and own the foundational infrastructure that our products run upon.
- Work directly on our products' golang code base to implement SRE related objectives.
- Take a data driven approach to quantifying system performance and reliability and use it to drive project priorities.
- Oncall participation including leading incident management for complex situations.
- Work on automation and advanced configuration management to allow our team to manage large numbers of clusters distributed across the world running various products.
- Work with infrastructure vendors when their solutions aren't meeting our real time performance and reliability needs.
Who You Are
- A balance of strengths in both software engineering and large scale system administration.
- Experience managing complex multi-region distributed systems running on top of container orchestration systems like Kubernetes.
- Passionate about maintainability and keeping system complexity at bay, but able to balance this with meeting launch deadlines.
Bonus Points
- Incident management training and experience being an Incident Commander.
- Experience with Linux networking, overlay networks, and Kubernetes CNIs.
- Low level knowledge for troubleshooting and tuning latency sensitive workloads.
Our Commitments to You
We offer:
- A competitive salary and equity package.
- Health, dental, and vision benefits
- Flexible vacations
- Remote work environment with necessary equipment provided.
Skills
Go, Kubernetes, Linux, SRE, Incident Management, Container Orchestration, Distributed Systems, Automation, Configuration Management
Similar jobs
DevOps / SRE jobsSenior Site Reliability Engineer responsible for designing and operating reliable, scalable production infrastructure, leading incident response, and improving observability and resilience. Requires 5+ years of reliability-focused engineering experience and expertise across cloud, infrastructure as code, Kubernetes, monitoring, and application development.
Leads end-to-end infrastructure for a scientific imaging platform, covering Linux administration, GPU/HPC systems, storage, upgrades, and vendor coordination. The role supports AI-enabled imaging workflows and requires extensive production Linux, Image Artist, GPU, HPC, and enterprise storage experience.
Leads Ireland-based Platform Developer Enablement and SRE teams, defining platform strategy, developer self-service, reliability objectives, and observability standards. Requires senior software, SRE, or platform engineering experience, management leadership, and expertise in cloud infrastructure, Kubernetes, Terraform, CI/CD, and distributed systems.
Senior network engineer responsible for designing, operating, and securing MongoDB’s global network and VPN infrastructure. The role requires 6+ years of networking or systems engineering experience, strong enterprise networking expertise, automation skills, and the ability to lead complex infrastructure initiatives.
Senior Site Reliability Engineer responsible for production troubleshooting, incident response, observability, SLOs, automation, and permanent reliability improvements. Requires strong software engineering, SQL, debugging, cloud-application troubleshooting, and cross-functional collaboration skills.