SRE / DevOps Lead
Leads a small SRE team while owning the reliability, security, scalability, and delivery infrastructure of a GCP-based platform. The role combines hands-on cloud operations with incident management, compliance support, observability, and corporate IT leadership.
About the job
Responsibilities
- Lead and grow a team of 2 SREs, setting direction across SRE, DevOps, and IT.
- Own the architecture, performance, and reliability of april's GCP-based platform, including Kubernetes, Terraform, CI/CD, networking, and scale.
- Define and uphold SLOs, infrastructure budgets, on-call priorities, and incident response practices.
- Build and operate the observability stack so engineers can identify and resolve problems proactively.
- Partner with engineering and security on controls, evidence, and audits supporting SOC 2, SOC 1, STIG, and IRS authorization compliance.
- Plan and execute capacity testing, load testing, and game days for tax season.
- Own corporate IT, including identity, Google Workspace, endpoint management, and the company SaaS stack.
- Establish standards for shipping software that is fast, safe, observable, and reversible.
Requirements
- 7+ years of SRE, DevOps, or platform engineering experience, including at least 1 year leading or managing engineers.
- Hands-on experience running production workloads on GCP, including Kubernetes (GKE) and Terraform at scale.
- Strong programming or scripting fundamentals in Go, Python, or a similar language, with a track record of automation.
- Experience operating in a SOC 2 environment and supporting external audits.
- Experience with regulated workloads such as IRS, financial services, or healthcare, or a demonstrated ability to learn regulated environments deeply.
- Experience building dashboards, alerts, and SLOs in Datadog or an equivalent observability platform.
- Calm, clear incident leadership and a coaching mindset.
- Working knowledge of identity and endpoint tools such as Okta, Google Workspace, Jamf, or Kandji.
- Clear written and spoken English and comfort collaborating across time zones.
Nice-to-haves
- Experience with the IRS e-file ecosystem, FTI handling, or fintech-grade data residency.
- Experience preparing infrastructure for SOC 2 Type II, ISO 27001, PCI, SOC 1, or STIG certification.
- FinOps or cloud cost management experience.
- Database operations at scale with Postgres, Spanner, or similar systems.
- Experience establishing or scaling an IT function at a high-growth startup.
- 3+ years of software engineering experience.
Skills
GCP, Kubernetes, Google Kubernetes Engine, Terraform, CI/CD, Go, Python, Datadog, Okta, Google Workspace, Jamf, Kandji, Postgres, Spanner
Similar jobs
DevOps / SRE jobsThe Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Owns and improves cloud infrastructure, CI/CD, Kubernetes, observability, security, scalability, and developer experience. The role requires at least six years of DevOps experience, strong AWS and automation expertise, and fluent Hebrew and English communication.
Designs and operates scalable, highly available cloud infrastructure while leading efficiency initiatives across compute, storage, networking, and cost optimization. Requires 5+ years of distributed-systems software development experience and expertise with cloud platforms, infrastructure as code, and Kubernetes.
Build and optimize ClickHouse Cloud’s highly available, multi-cloud infrastructure, including automation, distributed systems, networking, security, and cost-efficiency tooling. Requires 5+ years of experience operating scalable systems and expertise in cloud platforms, infrastructure as code, and production engineering.