Senior DevOps/SRE Engineer
Own reliability, deployments, observability, compliance, and AI infrastructure across AWS and Kubernetes for a fintech platform. The role requires strong DevOps/SRE depth, backend software engineering experience, and hands-on ownership of SOC 2 and PCI-DSS controls.
About the job
Responsibilities
Core reliability and infrastructure
- Own reliability, deployments, observability, and incident response across AWS, Kubernetes/EKS, Datadog, CI/CD, and networking.
- Maintain, optimize, and scale AWS infrastructure supporting backend, mobile, data, ML, and AI engineering teams.
- Own the infrastructure-as-code stack using Terraform/OpenTofu; AWS CDK experience is preferred.
- Favor open standards and swappable tooling to reduce vendor lock-in.
- Participate in an on-call rotation and handle interrupt-driven support.
- Improve efficiency and cost-effectiveness from investigation through implementation.
Compliance and security
- Own hands-on SOC 2 and PCI-DSS work by building and operating production systems that satisfy controls.
- Partner with security to close gaps, including network segmentation and traffic inspection.
CI/CD and deployments
- Make deployments and rollbacks easier, more predictable, and consistent.
- Consolidate and accelerate CI/CD pipelines supporting TypeScript/Node, Python, and Kotlin services.
AI infrastructure and product foundations
- Build internal AI infrastructure, including gateway and middleware layers, MCP integrations, knowledge graphs, and engineering tooling.
- Stand up greenfield infrastructure for new products without inheriting existing technical debt.
Requirements
- Strong DevOps/SRE experience and a genuine backend engineering background, including shipping production software.
- Hands-on AWS experience with observability and monitoring, CI/CD, compute, and networking.
- Production experience managing Kubernetes clusters in multiple environments.
- Experience with Terraform or comparable infrastructure-as-code tooling.
- Experience building DevOps/infrastructure at an early-stage startup or scaling it during significant growth.
- Willingness to own SOC 2 and PCI-DSS compliance work.
Nice-to-haves
- AI/LLM tooling, MCP, model gateways, or internal AI developer platforms.
- SOC 2 or PCI-DSS experience, including remediation of a specific control or finding.
- Security engineering exposure.
- AWS CDK, especially TypeScript.
- Reducing vendor lock-in or migrating data stores at scale.
Compensation and benefits
- Base salary: $170,000–$225,000 annually.
- Potential bonus and equity participation.
- Company-paid medical, dental, and vision coverage for employees and dependents from the first day.
- Up to $100 per month fitness reimbursement or complimentary LifeTime Fitness or Equinox membership.
- 401(k) with a 3.5% match and immediate vesting.
- Meal program for lunch and dinner.
- Pre-tax benefits, including a $1,000 HSA match.
- Life and accidental insurance.
- Flexible PTO.
- Employees are expected to work in the office five days per week, with flexibility for circumstances requiring remote work or adjusted schedules.
Skills
AWS, Kubernetes, Amazon Eks, Datadog, OpenTelemetry, CI/CD, Terraform, Opentofu, Aws Cdk, TypeScript, Node.js, Python, Kotlin, Mcp, SOC 2
Similar jobs
DevOps / SRE jobsOwn foundational cloud infrastructure and the internal developer platform supporting Commure’s engineering teams. The role requires 6+ years of infrastructure, platform, or SRE experience and hands-on expertise across Kubernetes, infrastructure as code, GitOps, observability, and cloud environments.
Leads cloud infrastructure, platform strategy, deployment pipelines, and infrastructure automation for a growing consumer platform. Requires 5+ years in infrastructure, DevOps, platform engineering, or SRE, plus deep AWS, coding, containerization, and infrastructure-as-code experience.
Own and evolve secure, highly available AWS and Azure infrastructure, including Terraform automation, Kubernetes, CI/CD, observability, networking, and incident response. The role requires 7+ years of DevOps or related experience and strong cross-functional partnership across engineering and security.
Own and evolve VSCO’s AWS/EKS platform, including infrastructure as code, GitOps, CI/CD, observability, networking, and production reliability. The role requires 5+ years of hands-on infrastructure or SRE experience and strong Kubernetes, Terraform, and AWS expertise.
Leads cross-functional technical initiatives and builds scalable business operations and customer-facing systems. Requires Python, system design, production engineering experience, and strong stakeholder collaboration; platform, AWS, SaaS, and analytics experience are preferred.