Production Engineer
Operates and improves Crusoe’s Kubernetes and virtual machine infrastructure, focusing on reliability, observability, incident response, and automation. Requires 3–6 years of production engineering experience, Kubernetes platform expertise, programming ability, and strong Linux and distributed-systems fundamentals.
About the job
Responsibilities
- Build and scale tooling and features for Crusoe’s Managed Kubernetes and Managed VM platforms for external customers.
- Collaborate on projects, incidents, priorities, and action plans for deploying new or retrofitting existing data centers.
- Advise software engineers on resilient coding practices and review changes before deployment.
- Review alerts and system performance metrics.
- Analyze system logs and develop monitoring enhancements.
- Participate in incident response drills, post-mortems, and root-cause analysis.
- Automate common error resolution through proactive remediation.
- Maintain high Service Level Indicators (SLIs) and Service Level Objectives (SLOs).
- Document work and share knowledge with the team.
Requirements
- 3–6 years of professional Production Engineering experience.
- Experience building Kubernetes platforms or Kubernetes controllers.
- Exposure to server-class hardware and provisioning.
- Understanding of distributed systems architecture, reliability, and scaling patterns.
- Basic understanding of infrastructure design and operational trade-offs involving networking, storage, and RPC serving.
- Proficiency in at least one programming language, such as Python or Go.
- Exposure to observability tooling and practices, including logging, monitoring, and alerting.
- Experience with Unix/Linux environments.
- Understanding of TCP/IP and network programming fundamentals.
- Awareness of information security best practices.
- Bachelor’s degree in Computer Science or a related field, or equivalent self-directed study of computer science fundamentals.
- Strong communication skills.
Benefits
- Pension contributions
- Private health and dental insurance
- Income protection
- Life assurance
- Competitive benefits supporting financial security, health, and well-being
Compensation
- Compensation is determined based on education, experience, knowledge, skills, abilities, internal equity, and market data.
Skills
Kubernetes, Kubernetes Controllers, Python, Go, Unix/Linux, Distributed Systems, TCP/IP, Network Programming, Observability, Logging, Monitoring, Alerting, Server Provisioning, Rpc
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.
Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.
Automate, manage, and optimize large-scale ClickHouse clusters handling trillions of events and 100+ PB data. Build provisioning systems with Terraform, Ansible, Kubernetes; focus on performance, scaling, and bleeding-edge features.
Build and operate Grafana’s physical infrastructure platform, including bare-metal environments, Kubernetes clusters, networking, scheduling, and autoscaling. The role requires datacenter and software-operations experience, with strong skills in Kubernetes and infrastructure automation using tools such as Go, Terraform, and Crossplane.