Infrastructure Engineer - Core Infrastructure
Operates and scales Kraken’s core infrastructure platforms, with a focus on OpenStack, Ceph, Linux, distributed systems, and automation. The role requires 3+ years of infrastructure or software engineering experience and supports reliable compute and storage services across cloud and on-premises environments.
About the job
Responsibilities
- Operate and evolve OpenStack platform components, including compute (Nova), networking (Neutron), and provisioning (Ironic).
- Learn and support storage systems, including Ceph architecture, operations, and troubleshooting.
- Handle feature requests and support tickets for OpenStack and storage systems.
- Troubleshoot complex issues across the infrastructure stack with senior engineers.
- Understand storage integration through OpenStack Cinder and Kubernetes CSI.
- Build automation and tooling to improve efficiency and reduce operational toil.
- Contribute to operational and troubleshooting documentation and runbooks.
- Participate in an on-call rotation to maintain platform reliability.
Requirements
- 3+ years of experience as an Infrastructure, Platform, DevOps, Software Engineer, or in a similar role.
- Experience building and maintaining an OpenStack private cloud or Ceph networked storage platform.
- Strong understanding of distributed systems fundamentals.
- Strong Linux systems knowledge, including shell usage, processes, networking basics, file systems, and permissions.
- Knowledge of networking fundamentals, including TCP/IP, DNS, ports, IP addressing, basic routing, TLS, and PKI concepts.
- Ability to use AI tools and agents such as Claude and OpenAI to deliver business value efficiently.
- Scripting or programming experience with Python, Bash, Go, or similar.
- Strong communication skills and a customer-focused approach to resolving compute and storage requests and enabling engineering teams.
Nice to Have
- Experience running Kubernetes clusters.
- Exposure to AWS, Google Cloud, Azure, or on-premises infrastructure.
- Knowledge of Terraform or other infrastructure-as-code tools.
- Experience with storage systems such as Ceph or Rook.
Compensation and Benefits
- Applications are accepted on an ongoing basis unless a specific deadline is stated in the posting.
Skills
Openstack, Ceph, Linux, Kubernetes, Python, Bash, Go, TCP/IP, DNS, Tls, Pki, Terraform, AWS, GCP, Azure
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Provides first-response incident triage and infrastructure stabilization for a production platform in a 24/7 rotation. Requires enterprise experience with Kubernetes, RabbitMQ, PostgreSQL, Azure, production troubleshooting, log-based diagnosis, and calm incident communication.
Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.
Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.