Senior Software Engineer - CI/CD
Builds and operates CI/CD platforms across GitHub Actions, Jenkins, Kubernetes, and cloud infrastructure. The role owns GitOps delivery, reusable developer tooling, observability, reliability, and migration initiatives across distributed engineering teams.
About the job
Responsibilities
- Design and operate a self-hosted GitHub Actions runner fleet on GKE, including autoscaling, reliability tuning, and zombie-runner cleanup.
- Own the ArgoCD topology powering CI/CD deployments, including central architecture, cluster connectivity, ApplicationSets, production reliability, and event-driven flows with Argo Events and GCP Pub/Sub.
- Maintain and evolve Helm charts and ApplicationSet patterns used by Kubernetes workloads, including versioning, release processes, backward compatibility, and developer experience.
- Own and evolve the Jenkins environment, including shared libraries, controllers and agents, plugins, credentials integration, JVM upgrades, and platform reliability.
- Build reusable GitHub Actions workflows, actions, and templates while driving migration from Jenkins and partnering on shared developer tooling.
- Lead platform projects from architectural design through implementation and long-term maintenance.
- Own pipeline reliability through Datadog observability, incident response, PagerDuty on-call rotations, secrets rotation, and vulnerability response.
- Partner with architects, infrastructure and security engineers, and product teams to develop golden paths, reusable workflows, and self-service tooling.
Requirements
- Deep expertise in Jenkins administration, Groovy shared libraries, Docker-based agents, GitHub Actions, reusable workflows, self-hosted runners, secrets management, and pipeline troubleshooting.
- Strong Kubernetes and GKE administration experience, including cluster optimization, networking, and Kubernetes internals.
- Hands-on experience with GCP and AWS services.
- Experience with ArgoCD, including ApplicationSets, sync strategies, and alerting.
- Experience authoring and versioning Helm charts.
- Experience with Argo Events or comparable event-driven delivery patterns.
- Strong proficiency with Terraform for GKE clusters, GCP resources, and CI/CD modules.
- Experience instrumenting and operating pipelines with Datadog, including custom metrics, dashboards, monitors, and distributed tracing.
- Production on-call experience with PagerDuty, including rotation management, alert routing, escalation policies, and post-incident reviews.
- Experience with container registries, including JFrog Artifactory, Xray/Curation scanning, GCP Artifact Registry, and image lifecycle policies.
- Production programming experience with at least one of Go, Python, or Node.js.
- Experience or knowledge building or deploying LLM-based applications, AI-assisted developer tooling, or AI infrastructure for engineering workflows.
Skills
Jenkins, Groovy, GitHub Actions, Kubernetes, Google Kubernetes Engine, GCP, AWS, Argo Cd, Helm, Terraform, Datadog, Pagerduty, Jfrog Artifactory, Docker
Similar jobs
DevOps / SRE jobsBuild and mature Mozilla’s internal developer infrastructure platform, including CI/CD, observability, Kubernetes optimization, environment bootstrapping, and cost optimization. The role requires 5+ years of software engineering experience, cloud-native expertise, and strong technical leadership.
Senior Software Engineer building and improving Mozilla’s internal developer infrastructure platform, including CI/CD, observability, Kubernetes, cloud optimization, and developer productivity workflows. Requires 5+ years of software engineering experience and expertise in cloud-native or platform engineering.
Build and operate scalable GPU/TPU HPC infrastructure for training and serving frontier AI models. The role partners with AI researchers, optimizes distributed workloads across clouds, and requires expertise in Kubernetes, Python, Go, Linux, and high-performance networking.
The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
Own and scale infrastructure for agent orchestration, sandboxing, and hosted MCP services. The role requires hands-on Kubernetes, cloud, and infrastructure-as-code experience, along with strong software engineering fundamentals and high ownership.