Senior Central Cloud Infrastructure Engineer
Senior engineer responsible for architecting and maintaining scalable AWS cloud infrastructure, leading modernization initiatives, and ensuring PCI/SOC2 compliance. Requires 5+ years experience with Terraform, Kubernetes, observability, and production cloud systems.
About the job
What you'll do
- Architect, implement, and maintain scalable, reliable, and secure cloud infrastructure environments across multiple company products
- Design and manage infrastructure modernization initiatives (CI/CD, IaC, K8s, AI, etc) and lead the complex integration of legacy or established systems into modern, cloud-native environments
- Craft comprehensive disaster recovery strategies and engineer production-grade technical solutions to safeguard system uptime
- Serve as a technical focal point for troubleshooting, system deep-dives, and resolving infrastructure bottlenecks across both Linux and Windows operating systems
- Ensure global cloud environments continuously meet and exceed stringent security guidelines and technical compliance frameworks, specifically PCI and SOC2
- Lead incident response and root-cause analysis, and champion blameless postmortems and error-budget discipline across the team
- Provide technical leadership and guidance to junior and mid-level engineers, helping to scale the infrastructure team's overall technical competency
What we're looking for
- Hold a BS degree in Computer Science (or a relevant technical/engineering discipline) and 5+ years of hands-on experience, or equivalent professional experience (8+ years) working in high-growth cloud environments
- Demonstrate deep production-level technical experience using AWS, Terraform, Atmos, and Python (or equivalent automation language - Bash/Go)
- Demonstrate hands-on expertise with a modern observability stack (Datadog or equivalent: APM/distributed tracing, logs, metrics, RUM), including defining SLOs/SLIs and managing monitors-as-code
- Leverage production experience with service mesh and L7 proxies (Envoy/Istio) and Kubernetes ingress in high-throughput request paths
- Utilize advanced operational knowledge of Kubernetes/EKS, Helm, and overall container orchestration
- Apply strong secrets-management and rotation practices (Infisical, Vault, or AWS Secrets Manager) in a PCI/SOC2 context
- Showcase hands-on experience working with managed streaming data infrastructure via AWS MSK (Kafka) alongside various SQL and NoSQL AWS managed database solutions, including Amazon Aurora
- Maintain a strong understanding of cloud networking architecture, security protocols, and firewalls
- Exhibit expert competency managing, configuring, and troubleshooting Linux (and Windows where applicable) server environments
- Use excellent written and verbal communication skills to translate complex infrastructure concepts into clear, actionable technical documentation
Nice to have
- Previous experience working inside innovative, high-growth tech start-ups or fast-paced corporate engineering environments
Compensation
- Anticipated base salary: $160,000 - $200,000 USD annually
- Access to healthcare benefits, 401(k) plan, short-term and long-term disability coverage, basic life insurance, stock option plan, bonus plans
Skills
AWS, Terraform, Python, Kubernetes, EKS, Helm, Datadog, Envoy, Istio, Kafka, Amazon Aurora, Linux, Windows, CI/CD, Iac
Similar jobs
DevOps / SRE jobsSenior software engineer responsible for operating and evolving Voltus’s infrastructure platform across AWS, Kubernetes, Nomad, observability, stateful systems, and developer tooling. The role requires 6+ years of engineering experience, deep production Kubernetes and AWS expertise, and strong Go or Python skills.
Own and evolve VSCO’s AWS/EKS platform, including infrastructure as code, GitOps, CI/CD, observability, networking, and production reliability. The role requires 5+ years of hands-on infrastructure or SRE experience and strong Kubernetes, Terraform, and AWS expertise.
Build and operate highly available, distributed platform services and cloud infrastructure for petabyte-scale observability products. The role requires 6+ years of experience, strong Java and AWS expertise, Kubernetes and Terraform production experience, and a bachelor’s degree or equivalent.
Own foundational cloud infrastructure and the internal developer platform supporting Commure’s engineering teams. The role requires 6+ years of infrastructure, platform, or SRE experience and hands-on expertise across Kubernetes, infrastructure as code, GitOps, observability, and cloud environments.
The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.