Software Engineer - Platform Metal
Build and operate Grafana’s physical infrastructure platform, including bare-metal environments, Kubernetes clusters, networking, scheduling, and autoscaling. The role requires datacenter and software-operations experience, with strong skills in Kubernetes and infrastructure automation using tools such as Go, Terraform, and Crossplane.
About the job
Responsibilities
- Manage the physical “Metal” environment from bare metal to Kubernetes.
- Manage cluster networking components, including load balancing, NAT, DNS, CNIs, cross-cluster communication, routing protocols, and architecture.
- Manage scheduling and autoscaling.
- Maintain Crossplane compositions and Terraform modules for cloud-provider resources common to users.
- Manage versioning and compatibility for Crossplane, Terraform Core, and providers.
- Work with Grafana Cloud application teams to understand their needs and prioritize platform capabilities.
- Participate in an on-call rotation.
Requirements
- Experience working with distributed systems.
- Experience with holistic software development, including design documentation, developer feedback, and integration testing.
- Experience operating software and supporting both developers and operators.
- Datacenter experience.
- Experience with Go, Python, and Shell.
- Comfortable operating in a remote-first, highly distributed company.
Nice-to-haves
- Experience with Cluster API, Tinkerbell, Talos, or Ceph.
- Experience working in or on open-source or community-based projects.
- Experience with AWS, Google Cloud, and Azure, including EKS, GKE, and AKS.
- Experience operating and managing Kubernetes workloads.
- Experience with Tanka and Jsonnet.
- Familiarity with Kubernetes scheduling and Karpenter.
- Experience with Terraform or Crossplane.
- Enjoyment of programming in Go and building tools, utilities, and exporters.
Compensation and benefits
- UK compensation range: £72,177–£86,612.
- Restricted Stock Units (RSUs) are included for all roles.
- 100% remote work with a global culture.
- 30 days of annual leave per year, including 3 Grafana Shutdown Days.
- In-person onboarding.
- Career growth opportunities and transparent communication.
Skills
Kubernetes, Go, Python, Shell, Terraform, Crossplane, Cluster Api, Tinkerbell, Talos, Ceph, AWS, GCP, Microsoft Azure, Jsonnet, Karpenter
Similar jobs
DevOps / SRE jobsBuild scalable software, automation, and frameworks for managing large AI network fabrics, including metrics, provisioning, monitoring, configuration, and remediation. The role requires deep networking expertise and a track record of designing reliable systems that orchestrate large device fleets.
Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Build and maintain Cloudflare’s deployment platform, enabling progressive rollouts, health-mediated releases, and automated workflows at scale. The role requires at least four years of software development experience, backend and frontend experience, and comfort with rapid delivery and on-call support.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Site Reliability Engineers build and operate reliable, scalable production infrastructure across GitLab’s Infrastructure Platforms teams. The role requires strong software engineering and operations fundamentals, Kubernetes and infrastructure-as-code experience, cloud expertise, and comfort with automation, observability, and incident response.