# Software Engineer, Infrastructure

**Company:** [Granica](https://hotfix.jobs/companies/granica)
**Location:** Remote
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** Go, Java, Python, Kubernetes, AWS, GCP, Azure, Docker, CI/CD, Terraform, Distributed Systems, Prometheus, Grafana, Spark, Trino
**Posted:** 2026-09-09

> Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

## Job Description

## Responsibilities
- Design and operate cloud infrastructure supporting petabyte- to exabyte-scale systems.
- Manage and evolve production Kubernetes clusters with high availability and predictable scaling.
- Build and optimize CI/CD pipelines for faster builds, reliable tests, and safer deployments.
- Improve developer experience by reducing friction and automating workflows.
- Contribute to end-to-end testing and release-confidence systems.
- Strengthen observability across logging, metrics, tracing, and alerting.
- Collaborate across engineering teams on architecture, reliability, and scaling challenges.
- Own customer-facing technical integration, guide deployments, troubleshoot infrastructure-level issues in customer environments, and ensure reliable operation across diverse cloud, data, and security stacks.
- Drive operational excellence through reliability, automation, and best practices.

## Requirements
- 5+ years of experience in infrastructure, platform, or distributed systems engineering.
- Strong coding skills in Go, Java, Python, or similar languages.
- Production-grade Kubernetes experience with strong cloud and containerization skills across AWS, Google Cloud, Azure, and Docker.
- Deep experience with CI/CD systems and cloud infrastructure.
- Proficiency with infrastructure as code, particularly Terraform.
- Ability to debug complex, cross-layer issues involving networking, storage, and runtime.
- Track record of designing scalable, reliable, and cost-efficient systems.
- Strong communication, ownership, and ability to thrive in a fast-paced startup environment.

## Nice-to-Haves
- Experience with data platforms or lakehouse systems.
- Familiarity with Spark, Trino, Iceberg, Delta, or Airflow.
- Exposure to observability stacks such as Prometheus, Grafana, ELK, or Datadog.

## Compensation & Benefits
- Competitive salary, meaningful equity, and performance bonus for top performers.
- 401(k) with company match.
- Comprehensive health coverage.
- Unlimited paid time off.
- Daily catered meals in the Mountain View office.
- Support for research, publication, and conference participation.

## Similar jobs

- [Software Engineer: Resiliency - Deploy at Scale](https://hotfix.jobs/jobs/21bb6f4f-daea-46ab-bac8-fff083d17451) - Cloudflare - London, United Kingdom
- [Release Engineer - Data Plane Internal Tooling and Productivity](https://hotfix.jobs/jobs/ee184847-0b76-42f4-8c1f-157f199626d3) - Clickhouse - Remote
- [Capacity Ops Engineer](https://hotfix.jobs/jobs/f1904714-7dd3-4ee3-9e7a-e4fcf52083bd) - Baseten - San Francisco, CA - $170k – $230k/yr
- [IT Security and Automation Engineer](https://hotfix.jobs/jobs/604b87b5-13a2-4bba-88b2-f7d0fbbad141) - Teleport - Remote - $149k – $258k/yr
- [Electrical Field Engineer - Data Center](https://hotfix.jobs/jobs/6bfa0e4c-9ccf-438a-b65e-cd4c6297762c) - Crusoe - Remote - $196k – $235k/yr

**Apply:** https://hotfix.jobs/jobs/585da47a-e02d-4c63-9816-248a2faa9b5b
**Canonical:** https://hotfix.jobs/jobs/585da47a-e02d-4c63-9816-248a2faa9b5b