Senior Software Engineer, Infrastructure
Own and evolve VSCO’s AWS/EKS platform, including infrastructure as code, GitOps, CI/CD, observability, networking, and production reliability. The role requires 5+ years of hands-on infrastructure or SRE experience and strong Kubernetes, Terraform, and AWS expertise.
About the job
Responsibilities
- Design, build, and operate AWS and EKS platforms using Terraform, Helm/Kustomize, and Flux GitOps.
- Own production cutovers end to end, including ingress, DNS, workers, cron jobs, telemetry, dashboards, and rollback plans.
- Partner with product and data teams on namespaces, secrets, preview environments, and shared infrastructure.
- Build and maintain GitHub Actions CI/CD, including self-hosted runners, image builds, and deployment pipelines.
- Support build and test pipelines, container images, and runners for Go, Python, PHP, Node, Java, Ruby, and similar ecosystems.
- Operate MySQL, Mongo, Valkey/Redis, and OpenSearch as platform services.
- Manage Cloudflare and CloudFront DNS, CDN behavior, WAF, and custom domains.
- Improve observability and autoscaling through monitors, dashboards, and external-metrics HPA.
- Participate in the infrastructure on-call rotation; troubleshoot Linux, Kubernetes, networking, and DNS issues and implement durable fixes.
- Review infrastructure pull requests and design documents, explain tradeoffs, and help teams consume the platform.
Requirements
- 5+ years of experience in infrastructure, platform, or site reliability engineering, including hands-on production ownership.
- Production experience with Kubernetes/EKS, Linux, and Terraform.
- Production AWS experience with EKS, IAM/IRSA, VPC, NAT, DNS, and at least one of RDS/MySQL, ElastiCache/Valkey, or OpenSearch.
- Experience with GitOps and CI/CD using Flux and/or Helm and GitHub Actions.
- Familiarity with Go, Python, PHP, Node, Java, Ruby, or similar languages in CI environments.
- Experience leading technical architecture discussions, explaining tradeoffs, and driving decisions.
- Experience troubleshooting complex production issues and converting them into durable fixes.
- Excellent written and verbal communication skills.
Nice to Have
- Cloudflare DNS, CDN, and zone migration experience.
- Experience with OpenSearch, Valkey/Redis, MySQL, Mongo, or Kafka.
- Experience with Datadog, Groundcover, Honeycomb, or OpenTelemetry.
- GitHub Actions Runner Controller, preview environments, or similar paved-road CI experience.
- Experience with Flux, Helm, and Kustomize.
- Experience using AI tools such as Cursor.
- Experience with high-scale consumer or creator products.
- Connection to photography, visual storytelling, or creative communities.
- Bachelor’s degree in Computer Science or a related technical discipline, or equivalent practical experience.
Compensation and Benefits
- Base salary: $165,000–$185,000 annually.
- Equity and discretionary performance bonuses.
- Flexible time off, 401(k), medical, dental, vision, life and disability insurance, paid holidays, and paid parental, medical, caregiver, and sick leave.
- Mental health resources and technology reimbursements.
Skills
AWS, Amazon Eks, Kubernetes, Terraform, Helm, Kustomize, Flux, GitOps, GitHub Actions, Linux, Cloudflare, Cloudfront, MySQL, MongoDB, Opensearch
Similar jobs
DevOps / SRE jobsOwn foundational cloud infrastructure and the internal developer platform supporting Commure’s engineering teams. The role requires 6+ years of infrastructure, platform, or SRE experience and hands-on expertise across Kubernetes, infrastructure as code, GitOps, observability, and cloud environments.
Leads cloud infrastructure, platform strategy, deployment pipelines, and infrastructure automation for a growing consumer platform. Requires 5+ years in infrastructure, DevOps, platform engineering, or SRE, plus deep AWS, coding, containerization, and infrastructure-as-code experience.
Senior software engineer responsible for operating and evolving Voltus’s infrastructure platform across AWS, Kubernetes, Nomad, observability, stateful systems, and developer tooling. The role requires 6+ years of engineering experience, deep production Kubernetes and AWS expertise, and strong Go or Python skills.
Own reliability, deployments, observability, compliance, and AI infrastructure across AWS and Kubernetes for a fintech platform. The role requires strong DevOps/SRE depth, backend software engineering experience, and hands-on ownership of SOC 2 and PCI-DSS controls.
Own and evolve secure, highly available AWS and Azure infrastructure, including Terraform automation, Kubernetes, CI/CD, observability, networking, and incident response. The role requires 7+ years of DevOps or related experience and strong cross-functional partnership across engineering and security.