Senior Dev Ops Engineer
Own and evolve secure, highly available AWS and Azure infrastructure, including Terraform automation, Kubernetes, CI/CD, observability, networking, and incident response. The role requires 7+ years of DevOps or related experience and strong cross-functional partnership across engineering and security.
About the job
Responsibilities
- Own and evolve cloud infrastructure across AWS and Azure.
- Design, implement, and maintain secure, scalable, highly available cloud architectures.
- Build and maintain infrastructure as code using Terraform and automation tools.
- Own and improve CI/CD pipelines, deployment processes, and release automation.
- Manage Kubernetes and containerized workloads.
- Establish observability across infrastructure and applications using monitoring, logging, alerting, and performance metrics with Datadog.
- Improve reliability, uptime, scalability, and disaster recovery.
- Partner with Engineering and Security to remediate infrastructure and cloud security risks.
- Manage cloud networking, identity management, object storage, relational databases, access controls, secrets management, DNS, and load balancing.
- Automate repetitive operational work and troubleshoot complex production issues.
- Lead infrastructure-related incident response.
- Establish infrastructure standards, documentation, and operational best practices.
- Support infrastructure requirements for high-scale AI/ML workloads.
- Mentor engineers and define the long-term cloud and DevOps roadmap.
- Support SOC 2, ISO, and FedRAMP certification efforts.
Requirements
- 7+ years of experience in DevOps, SRE, cloud infrastructure, or a related discipline.
- Expert-level experience with both AWS and Azure.
- Deep understanding of cloud architecture, networking, security, scalability, and reliability.
- Extensive hands-on experience with Terraform or comparable infrastructure-as-code tooling.
- Strong experience with Kubernetes and Docker/containerized environments.
- Experience building and managing CI/CD pipelines.
- Strong Linux systems administration and troubleshooting skills.
- Experience with VPCs/VNets, routing, DNS, load balancers, firewalls, and VPNs.
- Understanding of IAM, secrets management, encryption, least-privilege access, and cloud security.
- Experience with monitoring and observability platforms and practices.
- Strong scripting or programming ability in Python, Bash, Go, or similar languages.
- Experience operating production systems where reliability and uptime matter.
- Excellent incident response and troubleshooting skills.
- Ability to work independently, make sound technical decisions, and take ownership from architecture through operations.
- Strong communication and cross-functional collaboration skills.
Compensation and Benefits
- Salary: $170,000–$200,000 annually.
- Healthcare plans with 100% employee premium coverage and partial dependent coverage.
- Dental and vision plans with 100% employee premium coverage and dependent coverage.
- Short- and long-term disability and life insurance with 100% employee premium coverage.
- FSA/HSA and 401(k) programs.
- Equity compensation.
- 20 days of PTO per year.
- 12 weeks of parental leave.
- Learning and development budget.
- Monthly wellness benefits.
- Annual company-sponsored offsite.
- NYC HQ employees receive daily in-office lunch, commuter benefits, remote Fridays, happy hours, and other local events.
Skills
AWS, Azure, Terraform, Kubernetes, Docker, CI/CD, Linux, Datadog, Python, Bash, Go, Cloud Networking, IAM, DNS, Disaster Recovery
Similar jobs
DevOps / SRE jobsOwn foundational cloud infrastructure and the internal developer platform supporting Commure’s engineering teams. The role requires 6+ years of infrastructure, platform, or SRE experience and hands-on expertise across Kubernetes, infrastructure as code, GitOps, observability, and cloud environments.
Leads cloud infrastructure, platform strategy, deployment pipelines, and infrastructure automation for a growing consumer platform. Requires 5+ years in infrastructure, DevOps, platform engineering, or SRE, plus deep AWS, coding, containerization, and infrastructure-as-code experience.
Own reliability, deployments, observability, compliance, and AI infrastructure across AWS and Kubernetes for a fintech platform. The role requires strong DevOps/SRE depth, backend software engineering experience, and hands-on ownership of SOC 2 and PCI-DSS controls.
Own and evolve VSCO’s AWS/EKS platform, including infrastructure as code, GitOps, CI/CD, observability, networking, and production reliability. The role requires 5+ years of hands-on infrastructure or SRE experience and strong Kubernetes, Terraform, and AWS expertise.
Leads cross-functional technical initiatives and builds scalable business operations and customer-facing systems. Requires Python, system design, production engineering experience, and strong stakeholder collaboration; platform, AWS, SaaS, and analytics experience are preferred.