Senior Infra Engineer: Baremetal Orchestration
Builds and maintains bare metal provisioning, orchestration engine, and internal tools for Railway's infrastructure platform. Optimizes fleet efficiency and develops resilient services using Golang/Rust, Ansible, and Terraform for distributed systems.
About the job
Responsibilities
- Build and maintain host provisioning stack: PXE boot, Ansible, and burn-in agents for bare metal.
- Evolve homegrown orchestration engine for clusters, containers, and VMs.
- Optimize bin packing algorithm for utilization, performance, and cost minimization.
- Own internal tooling for Railway engineers interacting with the fleet.
- Build internal observability and alerting for fleet issues.
- Design and maintain CI pipelines for infrastructure code.
- Define immutable infrastructure using Terraform and Ansible.
- Build Golang/Rust gRPC services for millions of users.
- Write Engineering Requirement Documents from idea to implementation.
Requirements
- Strong understanding of distributed systems, fault tolerance, resilience, and scalability.
- Hands-on experience with bare metal provisioning, configuration management, and hardware production readiness.
- Comfort building and operating internal tools with focus on developer experience.
- Intuition for solution longevity (12-18 months in startups).
- Tact for implementation, monitoring error boundaries, and documentation.
- Strong prioritization in ambiguity, grit to scale and replace solutions.
- Excellent communication skills.
Nice-to-haves
- Experience with startups and high ownership culture.
Skills
Ansible, Terraform, Pxe, gRPC, Go, Rust, Kubernetes, Distributed Systems, CI/CD, Observability
Similar jobs
DevOps / SRE jobsDesigns and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.
Designs, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.
Build and operate core platform infrastructure, developer tooling, CI/CD, observability, and cloud reliability systems for a regulated payments platform. Requires 5+ years of infrastructure or backend experience, strong infrastructure-as-code skills, and production cloud expertise.
Build and mature Mozilla’s internal developer infrastructure platform, including CI/CD, observability, Kubernetes optimization, environment bootstrapping, and cost optimization. The role requires 5+ years of software engineering experience, cloud-native expertise, and strong technical leadership.
Senior Software Engineer building and improving Mozilla’s internal developer infrastructure platform, including CI/CD, observability, Kubernetes, cloud optimization, and developer productivity workflows. Requires 5+ years of software engineering experience and expertise in cloud-native or platform engineering.