Skip to content
AttainAttain

Sr/Staff Site Reliability Engineer, Consumer Apps

Sr/Staff SRE building and automating infrastructure for a fintech platform using AI agents, Terraform, Kubernetes, and GCP services. Focus on eliminating toil, improving observability, and ensuring reliability at scale for consumer apps processing billions in transactions.

About the job

Responsibilities

  • Use AI agents as a force multiplier for yourself and others.
  • Create, improve, and maintain internal agentic tools and harnesses.
  • Add automation to both existing and new systems until manual processes, and the toil that comes with them, simply go away.
  • Write Terraform modules for deploying infrastructure resources via our GitLab pipelines.
  • Develop Helm charts for deploying services and jobs in our Kubernetes cluster.
  • Define metrics, network policies, and routing rules for our Istio service mesh.
  • Monitor and maintain our GCP BigQuery, Spanner, and CloudSQL databases.
  • Pipe metrics to our Google-managed Prometheus instance and build out Grafana dashboards and alerts to increase visibility on our systems.
  • Experiment with GCP offerings, 3rd party vendors, AI tooling, and open-source projects to further automate and secure day-to-day operations.
  • Pair with engineering leads to instrument and monitor critical functionality.
  • Participate in architecture design and capacity planning discussions to ensure that our systems are scalable, maintainable, reliable, and secure.
  • Build, maintain, and improve our CI/CD pipeline.

Requirements

  • 6+ years of experience building and maintaining large-scale cloud-native infrastructure (AWS and/or GCP).
  • Demonstrated fluency directing AI coding agents (e.g. Claude Code, Cursor, or similar) to build, operate, and debug real infrastructure; and robust and experienced judgment on verification of their work.
  • A track record of replacing manual operations with durable automation.
  • Experience working with the containerization technologies Docker, Kubernetes, and Istio or a similar service mesh technology.
  • Experience with SQL database technologies such as MySQL, Google BigQuery, and Google Spanner.
  • Experience with stream technologies such as Kafka and Amazon Kinesis.
  • Experience with pub sub technologies such as AWS SNS and Google Pub/Sub.
  • Experience with serverless computing technologies such as AWS Lambda and Google Cloud Functions/Google Cloud Run.
  • Experience with infrastructure-as-code tools such as Terraform.
  • Experience with observability tools such as Datadog, Prometheus, and Grafana.
  • Strong computer science and software engineering fundamentals.
  • Experience with SOC2 and PCI Compliance processes and requirements.

Nice-to-Haves

  • You reach for automation before you reach for a runbook, and a manual process is something you want to delete, not document.
  • You treat AI agents as power tools and have real opinions about how to drive them — especially when to stop trusting them.
  • You are comfortable wearing many hats.
  • You have a willingness to learn and teach in a fast-paced, collaborative environment.
  • You have a strong desire to automate things.
  • You readily provide constructive feedback, and also proactively seek feedback to improve yourself.
  • You like to get your hands dirty and tinker with/stress test new technologies.

Skills

Terraform, Kubernetes, Istio, Docker, GCP, BigQuery, Spanner, Cloudsql, Prometheus, Grafana, AI Agents, Kafka, AWS Lambda, Cloud Run

Anthropic

Anthropic

San Francisco, CA
Staff+ Site Reliability Engineer, Safeguards ML Infra
$320k+/yrHybrid8+ YOEDevOps / SRE

Staff-level site reliability engineer responsible for safely deploying and operating safeguards infrastructure across model releases and cloud platforms. The role emphasizes production change management, high-stakes incident response, and automating manual launch and validation processes.

Polymarket

Polymarket

New York, NY

Staff Infrastructure Engineer
$250k+/yrOn-site7+ YOEDevOps / SRE

Staff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.

Fortanix

Fortanix

Santa Clara, CA

Senior/Staff Infrastructure & Platform Engineer
$155k+/yrOn-site7+ YOEDevOps / SRE

Leads the architecture, development, and operation of cloud, Kubernetes, on-premises, and hybrid infrastructure, while building developer platforms and CI/CD automation. Requires at least six years of infrastructure or related engineering experience, deep Kubernetes expertise, strong programming skills, and technical leadership.

Scale AI

Scale AI

San Francisco, CA

Staff Network Engineer, App Platform
No salary listedOn-site7+ YOEDevOps / SRE

Own the network architecture and standards for a multi-cloud enterprise AI platform deployed across Kubernetes environments and customer-controlled networks. The role requires deep cloud and Kubernetes networking expertise, strong security fundamentals, and the judgment to establish scalable, supportable connectivity patterns.

Motive

Motive

Buffalo, NY
Staff Platform Engineer
$164k+/yrOn-site7+ YOEDevOps / SRE

Staff Platform Engineer will build and improve automated delivery pipelines, developer environments, infrastructure, and release systems across the engineering organization. The role requires 6+ years of engineering experience, a bachelor’s degree, and expertise with CI/CD, cloud infrastructure, containers, and infrastructure as code.