Senior Cloud Engineer
Senior Cloud Engineer owning AWS/GCP infrastructure, Kubernetes/GitOps platforms, and CI/CD systems. Designs and operates scalable, secure cloud infrastructure while mentoring engineers and enabling AI/ML tooling.
About the job
Responsibilities
- Design, build, and operate cloud infrastructure, platform tooling, and CI/CD systems for OfferUp's core applications
- Build cloud infrastructure as code using Terraform on AWS and GCP, optimizing for reliability, security, and cost
- Own and evolve the Kubernetes & GitOps platform (EKS/GKE, Envoy/Gloo networking, Helm, ArgoCD)
- Build and maintain CI/CD pipelines and developer workflows (GitHub Actions, monorepo tooling, JFrog Artifactory)
- Maintain backend "golden templates" and shared libraries used across the organization
- Lead platform migrations end-to-end including observability, data-store, Kubernetes, and multi-account migrations
- Drive operational excellence and security: own observability in Datadog, participate in on-call, remediate CVEs, manage secrets and access controls (Cloudflare/ZTNA, Okta, SSO, Workload Identity Federation), optimize cloud cost
- Enable AI/ML developer tooling (Amazon Bedrock, Claude, AI-assisted workflows)
- Lead design and code reviews, set technical standards, mentor engineers, and partner with other teams
Requirements
- 5–8 years in cloud engineering, infrastructure, platform, DevOps, or SRE roles
- Bachelor's in Computer Science or a related field, or equivalent practical experience
- Deep, hands-on experience with AWS (ideally GCP): EKS/GKE, Lambda, networking, IAM, S3, Kinesis, OpenSearch, Route53 using Terraform at scale
- Solid experience with Kubernetes (EKS/GKE), Helm, GitOps (ArgoCD), and CI/CD pipelines (GitHub Actions or similar) in a monorepo or multi-service environment
- Experience with modern observability tooling (Datadog or similar) for metrics, logs, traces, and alerting; supporting production systems, on-call, and debugging complex distributed systems
- Proficiency in at least one of Java, TypeScript, Python, or Go for automation, tooling, and platform development; comfort with scripting and SQL
- Ability to communicate technical concepts to technical and non-technical audiences, lead through influence, and mentor other engineers
Nice-to-Haves
- Experience with edge and zero-trust networking (Cloudflare, Envoy/Gloo)
- Experience operating data and streaming infrastructure (Kinesis, Flink, Redis/Valkey, Confluent/Kafka)
- Experience enabling AI/ML infrastructure or developer tooling (Amazon Bedrock, SageMaker, or LLM-based developer workflows)
- Experience with security and compliance practices, change management, and incident management at scale
- Experience with GCP Workload Identity Federation and cross-cloud (AWS↔GCP) integrations
Compensation & Benefits
- Compensation Range: $215,000 - $240,000
- Equity in OfferUp
- Health insurance, healthcare savings and spending accounts
- 401(k) plan with match
- Basic and voluntary life insurance, disability benefits
- Paid time off: sick leave, family/medical leave, vacation (flexible 3-5 weeks), 12 company holidays
Skills
AWS, GCP, Terraform, Kubernetes, EKS, GKE, Helm, Argo CD, GitHub Actions, Datadog, Java, TypeScript, Python, Go, CI/CD
Similar jobs
DevOps / SRE jobsSenior engineer owning safety-critical software pipelines and infrastructure, from static and dynamic analysis through CI enforcement, dashboards, and reliability tooling. Requires an advanced technical degree, 7+ years working with large codebases, and expertise in Bazel, Python, backend infrastructure, and C++.
Own and evolve a broad infrastructure platform spanning cloud, Kubernetes, deployment, reliability, security, and GPU-backed AI systems. The role requires 8+ years operating production distributed systems, strong incident and architecture experience, and practical cloud infrastructure expertise.
Build and evolve the developer platform that enables reliable, efficient software delivery across the company. The role requires 5+ years of software engineering experience, strong programming and system-design fundamentals, and expertise in build systems, CI/CD, testing, and deployment automation.
Build and improve cloud infrastructure, developer workflows, and internal tooling that make software development, testing, and releases more efficient and reliable. The role requires cloud architecture knowledge, CI/CD experience, Terraform and Bazel proficiency, and software development skills in Go, Python, or C++.
Build and operate scalable control-plane and data-plane infrastructure for distributed AI workloads, including Ray cluster orchestration, scheduling, observability, and accelerator integration. Requires a bachelor's degree or equivalent experience, 3+ years of production coding, cloud-native expertise, Kubernetes, and Go/Python proficiency.