Software Engineer, Infrastructure
Owns and evolves Kubernetes-based infrastructure for secure, compliant AI deployments in financial services, including observability with Datadog, IaC with Terraform, and incident response. Requires 8+ years experience with Docker, K8s, AWS, and Python at scale.
About the job
What You’ll Do
- Own and evolve our Kubernetes infrastructure, including cluster management, service mesh configuration, and container security policies.
- Design and implement progressive delivery pipelines with canary deployments, automated rollbacks, and deployment health validation.
- Build and maintain our observability infrastructure in Datadog, including dashboards, monitors, SLOs, and distributed tracing.
- Drive incident response for high-severity outages and proactively model capacity needs for low-latency AI inference.
- Architect and automate secure infrastructure using Infrastructure-as-Code for VPCs, IAM policies, Kubernetes manifests, and private cloud deployments.
- Maintain and improve the infrastructure controls that support our SOC 2 compliance posture.
- Lead customer engagements for enterprise rollouts and mentor mid-level engineers on infrastructure best practices.
What We’re Looking For
Must-Haves:
- 8+ years in infrastructure engineering or DevOps at high-growth or hyperscale companies.
- Experience with Docker and Kubernetes, including production cluster management, Helm, and service mesh technologies.
- A proven track record of architecting and operating AWS (preferred), GCP, or Azure at an enterprise scale.
- Experience with observability platforms, preferably Datadog (metrics, logs, APM, distributed tracing).
- A strong background in Infrastructure-as-Code (Terraform, Helm, Kustomize) and safe deployment practices (progressive delivery, canary deployments, GitOps, automated rollbacks).
- "Battle scars" from leading outages, capacity events, and large-scale incident reviews.
- Strong programming skills in Python.
Bonus Points:
- Familiarity with TypeScript.
- Direct involvement in SOC 2 or other compliance audit preparation or remediation.
- Direct experience with private-cloud or on-premises deployments for regulated customers.
- Previous experience at startups scaling infrastructure from the early stages to the enterprise level.
- A background in fintech or building systems for highly regulated industries.
- Experience with AI/ML infrastructure and model deployment at scale.
Compensation & Benefits
- $168k - $213k + equity
- Comprehensive healthcare, 401k matching, commuter benefits
- 15 days PTO + holidays, unlimited sick days
- Flexible leave options
Skills
Kubernetes, Docker, Helm, Datadog, Terraform, AWS, Python, GitOps, Service Mesh, Kustomize
Similar jobs
DevOps / SRE jobsOwn foundational cloud infrastructure and the internal developer platform supporting Commure’s engineering teams. The role requires 6+ years of infrastructure, platform, or SRE experience and hands-on expertise across Kubernetes, infrastructure as code, GitOps, observability, and cloud environments.
Leads cloud infrastructure, platform strategy, deployment pipelines, and infrastructure automation for a growing consumer platform. Requires 5+ years in infrastructure, DevOps, platform engineering, or SRE, plus deep AWS, coding, containerization, and infrastructure-as-code experience.
Own reliability, deployments, observability, compliance, and AI infrastructure across AWS and Kubernetes for a fintech platform. The role requires strong DevOps/SRE depth, backend software engineering experience, and hands-on ownership of SOC 2 and PCI-DSS controls.
Own and evolve secure, highly available AWS and Azure infrastructure, including Terraform automation, Kubernetes, CI/CD, observability, networking, and incident response. The role requires 7+ years of DevOps or related experience and strong cross-functional partnership across engineering and security.
Own and evolve VSCO’s AWS/EKS platform, including infrastructure as code, GitOps, CI/CD, observability, networking, and production reliability. The role requires 5+ years of hands-on infrastructure or SRE experience and strong Kubernetes, Terraform, and AWS expertise.