DevOps Engineer, Cloud Platform
Build and operate shared Kubernetes (EKS) and AWS cloud infrastructure powering Upstart's product and ML workloads. Requires 3+ years Kubernetes production experience plus strong AWS, IaC, and GitOps skills.
About the job
How you’ll make an impact
- Design and operate a fleet of Kubernetes (EKS) clusters across production, staging, and ephemeral environments, ensuring reliability and high availability
- Evolve AWS infrastructure and network architecture (VPCs, subnets, IAM, account structure) to support scalable, multi-team workloads
- Build and maintain infrastructure-as-code and GitOps workflows using tools such as Terraform, CDK, and ArgoCD
- Improve platform reliability and performance by defining and driving SLOs, analyzing incidents, and implementing systemic fixes
- Participate in and help improve the on-call rotation, leading incident response and post-incident reviews to drive systemic platform improvements
- Partner with SRE, Delivery, InfoSec, and product/ML teams to land high-impact infrastructure changes and platform standards
- Drive improvements in developer experience by simplifying platform usage, reducing toil, and enabling faster product and ML development
- Contribute to cost efficiency initiatives by optimizing resource utilization across Kubernetes and cloud infrastructure
Minimum Qualifications
- Bachelor’s degree in Computer Science, Engineering, Mathematics, or a related field (or equivalent practical experience) and 3+ years of professional experience
- 3+ years of experience operating Kubernetes in production environments, including cluster networking, storage, and RBAC
- Proficiency with AWS infrastructure, including VPC design, networking, and IAM
- Proven expertise in implementing infrastructure-as-code using tools such as Terraform or AWS CDK
- Experience implementing GitOps workflows using tools such as ArgoCD or similar
- Ability to influence technical decisions across teams and drive adoption of platform standards
Preferred Qualifications
- Knowledge of service mesh technologies such as Istio or Envoy
- Experience designing or operating multi-cluster Kubernetes architectures
- Experience with cloud networking at scale, including ingress/egress or edge platforms (e.g., Cloudflare)
- Knowledge of cloud security, identity, and compliance frameworks (e.g., IAM, SOC 2, CIS benchmarks)
Skills
Kubernetes, AWS, Terraform, Aws Cdk, Argo CD, EKS, GitOps, Istio, Envoy, SLOs, IAM, Vpc
Similar jobs
DevOps / SRE jobsBuild and operate deployment platforms, automation, and developer tooling that make software releases safer, more reliable, and self-service. The role requires a bachelor’s degree or equivalent, three years of software engineering experience, and experience with production systems and cloud or distributed infrastructure.
Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.
Build and operate highly available infrastructure for an enterprise AI platform, spanning cloud systems, Kubernetes, automation, observability, and reliability engineering. Requires 5+ years of production infrastructure experience, strong Python or Go skills, and daily use of AI-assisted workflows.
Builds and scales highly available infrastructure using AWS, Terraform, and Docker to support rapid growth and AI workloads. Collaborates with product and research teams on architectures, CI/CD, monitoring, and performance optimization.
Build and operate Mercor’s enterprise agent platform across security, routing, isolated execution, orchestration, deployment, and production scalability. The role requires 5+ years building high-scale platforms, architectural ownership, and experience with core infrastructure primitives across multiple clouds.