Senior Software Engineer – Platform & Data Infrastructure
Senior engineer responsible for scaling and operating multi-region Kubernetes, GitOps, Infrastructure as Code, security governance, and data-platform infrastructure. The role requires 8+ years of platform, SRE, or cloud data infrastructure experience and strong Kubernetes and Terraform expertise.
About the job
Responsibilities
- Operate production, multi-region Kubernetes fleets, including lifecycle management, network topology, node-pool migrations, and deployments on GKE and EKS.
- Maintain automated GitOps reconciliation pipelines using FluxCD and Helm, including cluster bootstrapping, progressive rollouts, and secret integrations.
- Architect modular Terraform codebases, manage enterprise workspaces with Terraform Enterprise and Terragrunt, and enforce compliance guardrails using Sentinel and OPA policy-as-code.
- Build and optimize data pipelines and workflow orchestration using Apache Airflow, Kafka, and enterprise data warehouses such as BigQuery.
- Integrate enterprise secrets-management tools including HashiCorp Vault, External Secrets Operator, and Google Secret Manager/KMS.
- Enforce Workload Identity and network perimeter controls.
- Lead zero-downtime infrastructure initiatives, including IP-addressing migrations, cluster rebuilds, and database cutovers.
- Participate in on-call rotations, design observability dashboards with Splunk and GCP Cloud Logging/Stackdriver, write runbooks, and execute regional disaster-recovery failover drills.
- Mentor mid-level and junior engineers through code and design reviews and establish self-service infrastructure standards.
Requirements
- Bachelor’s or graduate degree in Computer Science, Software Engineering, or a related technical field.
- 8+ years of professional experience in platform engineering, Site Reliability Engineering, or cloud data infrastructure.
- Strong programming skills in Python, Bash, SQL, or Java.
- Deep experience operating enterprise-scale Kubernetes fleets and container ecosystems, particularly GKE or EKS.
- Extensive practical experience writing modular Terraform and managing Infrastructure as Code at scale.
- Experience operating event-streaming and workflow systems such as Apache Airflow, Kafka, or BigQuery.
Nice-to-haves
- Certified Kubernetes Administrator (CKA) or Certified Kubernetes Application Developer (CKAD) certification.
- Experience with Terraform Enterprise and Sentinel or Open Policy Agent policy-as-code guardrails.
- Experience with GitOps workflows using FluxCD or ArgoCD across multi-tenant production fleets.
- Advanced secrets-management expertise with HashiCorp Vault and Kubernetes External Secrets Operator.
- Experience executing large-scale cloud migrations, data warehouse transitions, or networking range overhauls with zero customer impact.
Compensation and Benefits
- Annual base salary: $190,978–$214,052 USD.
- Base salary excludes company bonus, sales incentives, equity, and benefits.
- Benefits include medical, dental, and vision coverage; health savings and flexible spending accounts; life and disability insurance; 401(k) with company match; parental leave; paid time off and company holidays; commuter benefits; learning and development support; and additional wellbeing and employee assistance programs.
Skills
Kubernetes, GKE, EKS, Terraform, GitOps, Fluxcd, Helm, Apache Airflow, Kafka, BigQuery, Python, Bash, SQL, Hashicorp Vault, Opa
Similar jobs
DevOps / SRE jobsOwn the reliability, resilience, observability, and automation of AWS and Kubernetes infrastructure supporting production products and AI/ML workloads. The role requires 4+ years of cloud infrastructure experience, strong Kubernetes and Terraform expertise, and senior-level incident response and software engineering skills.
Build and operate core platform infrastructure, developer tooling, CI/CD, observability, and cloud reliability systems for a regulated payments platform. Requires 5+ years of infrastructure or backend experience, strong infrastructure-as-code skills, and production cloud expertise.
Build and operate multi-cloud, multi-cluster infrastructure and platform primitives for large-scale simulations and enterprise AI workloads. The role requires 5+ years in infrastructure, platform, SRE, or DevOps systems, strong Kubernetes and cloud expertise, production programming skills, and Infrastructure as Code experience.
Own and evolve AWS cloud infrastructure, deployment, reliability, observability, and security for a growing financial and hospitality technology platform. The hands-on role requires 8+ years operating production cloud infrastructure, strong AWS and container orchestration expertise, and experience with migrations and incident response.
Senior SRE who embeds with product teams to improve reliability, observability, performance, and incident preparedness. The role requires SRE or DevOps experience, strong PostgreSQL and Temporal expertise, and familiarity with observability platforms and OpenTelemetry.