Sr/Staff Site Reliability Engineer, Consumer Apps
Sr/Staff SRE building and automating infrastructure for a fintech platform using AI agents, Terraform, Kubernetes, and GCP services. Focus on eliminating toil, improving observability, and ensuring reliability at scale for consumer apps processing billions in transactions.
About the job
Responsibilities
- Use AI agents as a force multiplier for yourself and others.
- Create, improve, and maintain internal agentic tools and harnesses.
- Add automation to both existing and new systems until manual processes, and the toil that comes with them, simply go away.
- Write Terraform modules for deploying infrastructure resources via our GitLab pipelines.
- Develop Helm charts for deploying services and jobs in our Kubernetes cluster.
- Define metrics, network policies, and routing rules for our Istio service mesh.
- Monitor and maintain our GCP BigQuery, Spanner, and CloudSQL databases.
- Pipe metrics to our Google-managed Prometheus instance and build out Grafana dashboards and alerts to increase visibility on our systems.
- Experiment with GCP offerings, 3rd party vendors, AI tooling, and open-source projects to further automate and secure day-to-day operations.
- Pair with engineering leads to instrument and monitor critical functionality.
- Participate in architecture design and capacity planning discussions to ensure that our systems are scalable, maintainable, reliable, and secure.
- Build, maintain, and improve our CI/CD pipeline.
Requirements
- 6+ years of experience building and maintaining large-scale cloud-native infrastructure (AWS and/or GCP).
- Demonstrated fluency directing AI coding agents (e.g. Claude Code, Cursor, or similar) to build, operate, and debug real infrastructure; and robust and experienced judgment on verification of their work.
- A track record of replacing manual operations with durable automation.
- Experience working with the containerization technologies Docker, Kubernetes, and Istio or a similar service mesh technology.
- Experience with SQL database technologies such as MySQL, Google BigQuery, and Google Spanner.
- Experience with stream technologies such as Kafka and Amazon Kinesis.
- Experience with pub sub technologies such as AWS SNS and Google Pub/Sub.
- Experience with serverless computing technologies such as AWS Lambda and Google Cloud Functions/Google Cloud Run.
- Experience with infrastructure-as-code tools such as Terraform.
- Experience with observability tools such as Datadog, Prometheus, and Grafana.
- Strong computer science and software engineering fundamentals.
- Experience with SOC2 and PCI Compliance processes and requirements.
Nice-to-Haves
- You reach for automation before you reach for a runbook, and a manual process is something you want to delete, not document.
- You treat AI agents as power tools and have real opinions about how to drive them — especially when to stop trusting them.
- You are comfortable wearing many hats.
- You have a willingness to learn and teach in a fast-paced, collaborative environment.
- You have a strong desire to automate things.
- You readily provide constructive feedback, and also proactively seek feedback to improve yourself.
- You like to get your hands dirty and tinker with/stress test new technologies.
Skills
Terraform, Kubernetes, Istio, Docker, GCP, BigQuery, Spanner, Cloudsql, Prometheus, Grafana, AI Agents, Kafka, AWS Lambda, Cloud Run
Similar jobs
DevOps / SRE jobsStaff-level site reliability engineer responsible for safely deploying and operating safeguards infrastructure across model releases and cloud platforms. The role emphasizes production change management, high-stakes incident response, and automating manual launch and validation processes.
Staff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.
Leads the architecture, development, and operation of cloud, Kubernetes, on-premises, and hybrid infrastructure, while building developer platforms and CI/CD automation. Requires at least six years of infrastructure or related engineering experience, deep Kubernetes expertise, strong programming skills, and technical leadership.
Own the network architecture and standards for a multi-cloud enterprise AI platform deployed across Kubernetes environments and customer-controlled networks. The role requires deep cloud and Kubernetes networking expertise, strong security fundamentals, and the judgment to establish scalable, supportable connectivity patterns.
Staff Platform Engineer will build and improve automated delivery pipelines, developer environments, infrastructure, and release systems across the engineering organization. The role requires 6+ years of engineering experience, a bachelor’s degree, and expertise with CI/CD, cloud infrastructure, containers, and infrastructure as code.