Skip to content
KrakenKraken

Infrastructure Engineer - Core Infrastructure

Operates and scales Kraken’s core infrastructure platforms, with a focus on OpenStack, Ceph, Linux, distributed systems, and automation. The role requires 3+ years of infrastructure or software engineering experience and supports reliable compute and storage services across cloud and on-premises environments.

About the job

Responsibilities

  • Operate and evolve OpenStack platform components, including compute (Nova), networking (Neutron), and provisioning (Ironic).
  • Learn and support storage systems, including Ceph architecture, operations, and troubleshooting.
  • Handle feature requests and support tickets for OpenStack and storage systems.
  • Troubleshoot complex issues across the infrastructure stack with senior engineers.
  • Understand storage integration through OpenStack Cinder and Kubernetes CSI.
  • Build automation and tooling to improve efficiency and reduce operational toil.
  • Contribute to operational and troubleshooting documentation and runbooks.
  • Participate in an on-call rotation to maintain platform reliability.

Requirements

  • 3+ years of experience as an Infrastructure, Platform, DevOps, Software Engineer, or in a similar role.
  • Experience building and maintaining an OpenStack private cloud or Ceph networked storage platform.
  • Strong understanding of distributed systems fundamentals.
  • Strong Linux systems knowledge, including shell usage, processes, networking basics, file systems, and permissions.
  • Knowledge of networking fundamentals, including TCP/IP, DNS, ports, IP addressing, basic routing, TLS, and PKI concepts.
  • Ability to use AI tools and agents such as Claude and OpenAI to deliver business value efficiently.
  • Scripting or programming experience with Python, Bash, Go, or similar.
  • Strong communication skills and a customer-focused approach to resolving compute and storage requests and enabling engineering teams.

Nice to Have

  • Experience running Kubernetes clusters.
  • Exposure to AWS, Google Cloud, Azure, or on-premises infrastructure.
  • Knowledge of Terraform or other infrastructure-as-code tools.
  • Experience with storage systems such as Ceph or Rook.

Compensation and Benefits

  • Applications are accepted on an ongoing basis unless a specific deadline is stated in the posting.

Skills

Openstack, Ceph, Linux, Kubernetes, Python, Bash, Go, TCP/IP, DNS, Tls, Pki, Terraform, AWS, GCP, Azure

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

Perplexity

Perplexity

San Francisco, CA
Member of Technical Staff
$220k+/yrRemote4+ YOEDevOps / SRE

Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.

Invisible Tech

Invisible Tech

Estonia
Site Reliability Engineer
No salary listedRemoteDevOps / SRE

Provides first-response incident triage and infrastructure stabilization for a production platform in a 24/7 rotation. Requires enterprise experience with Kubernetes, RabbitMQ, PostgreSQL, Azure, production troubleshooting, log-based diagnosis, and calm incident communication.

Supabase

Supabase

Remote

Platform Engineer - Compute Capacity
No salary listedRemote5+ YOEDevOps / SRE

Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.

Alpaca

Alpaca

Remote

Production Support Engineer
No salary listedRemote4+ YOEDevOps / SRE

Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.